|
Mingkai Chen
Email  / 
CV  / 
Google Scholar  / 
GitHub  / 
LinkedIn
I am a Member of Technical Staff at Gyges Labs in San Mateo, California,
where I build the AI software behind Vocci, an AI note-taking ring.
My work covers the assistant's memory architecture, speaker recognition, the language-model pipeline,
and making these models run efficiently on a small wearable device.
Before joining Gyges Labs full time in December 2025, I was a Ph.D. student in
Electrical and Computer Engineering at the
Rochester Institute of Technology (2024–2025), advised by
Prof. Dongfang Liu, working on large language models,
multimodal models, and generative models for scientific data.
I received my B.S. in Computer Science from
Stony Brook University in 2024.
|
|
|
Work
Gyges Labs Inc, San Mateo, CA — Member of Technical Staff (Dec 2025–present);
Research Intern (May–Dec 2025).
At Gyges Labs I research, develop and evaluate the AI algorithms for the company's wearable devices:
the native memory architecture that gives the assistant persistent,
context-aware memory across sessions; multimodal integration and continual learning over the audio
and text the device captures; on-device and server optimisation, benchmarking and reliability testing
ahead of the Vocci launch (Aug 2026); and
Vocci MCP, the Model Context Protocol integration that passes recorded
conversation context to third-party AI tools. During my internship I built the streaming speaker-recognition
system, the first version of the memory module, and re-architected the AI serving infrastructure.
|
|
Research
My research interests are in
large language models,
multimodal models, and
generative models for scientific data.
All of my papers are listed below with links.
|
Diff-PIC: Revolutionizing Particle-In-Cell Nuclear Fusion Simulation with Diffusion Models
Chuan Liu,
Chunshu Wu,
Shihui Cao,
Mingkai Chen,
James Chenhao Liang,
Ang Li,
Michael Huang,
Chuang Ren,
Ying Nian Wu,
Dongfang Liu,
Tong Geng
ICLR 2025
arXiv
We developed Diff-PIC, a paradigm using conditional diffusion models to efficiently simulate Laser-Plasma Interaction (LPI) for nuclear fusion research. Designed a distillation process to capture physical patterns from Particle-in-Cell (PIC) simulations, and addressed key challenges with a physically-informed model and rectified flow technique, enhancing efficiency and fidelity. This innovation significantly reduces computational barriers in nuclear fusion research, advancing sustainable energy solutions.
|
Inertial Confinement Fusion Forecasting via Large Language Models
Mingkai Chen,
Taowen Wang,
Shihui Cao,
James Chenhao Liang,
Chuan Liu,
Chunshu Wu,
Qifan Wang,
Ying Nian Wu,
Michael Huang,
Chuang Ren,
Ang Li,
Tong Geng,
Dongfang Liu
arXiv preprint, 2024; under journal review
We developed LPI-LLM, integrating Large Language Models with reservoir computing to address Laser-Plasma Instabilities (LPI) in Inertial Confinement Fusion (ICF). Our designed fusion-specific models for accurate hot electron predictions and actionable insights, achieving state-of-the-art performance in forecasting Hard X-ray (HXR) energies with a negligible computational cost compared to traditional simulation methods. In addition, we created LPI4AI, the first experimental benchmark for advancing AI-driven fusion research.
|
A Benchmark and Chain-of-Thought Prompting Strategy for Large Multimodal Models with Multiple Image Inputs
Daoan Zhang,
Junming Yang,
Hanjia Lyu,
Zijian Jin,
Yuan Yao,
Mingkai Chen,
Jiebo Luo
ICPR 2024 (Lecture Notes in Computer Science, vol. 15318, pp. 226–241)
arXiv
We investigated Large Multimodal Models' (LMMs) ability to process multiple image inputs, focusing on fine-grained perception and information blending. Our research involved image-to-image matching and multi-image-to-text matching assessments, using models like GPT-4V and Gemini. We developed a Contrastive Chain-of-Thought (CoCoT) prompting method to improve LMMs' multi-image understanding, significantly enhancing model performance in our evaluations.
|
Aggregation of Disentanglement: Reconsidering Domain Variations in Domain Generalization
Daoan Zhang*,
Mingkai Chen*,
Chenming Li,
Lingyun Huang,
Jianguo Zhang
arXiv preprint, 2023; under review, International Journal of Computer Vision
We proposed a new perspective to utilize class-aware domain variant features in training, and in the
inference period, our model effectively maps target domains into the latent space where the known
domains lie. We also designed a contrastive learning based paradigm to calculate the weights for
unseen domains.
|
|
* equal contribution.
|
|
Service
Reviewer for Neurocomputing (Elsevier), IEEE Transactions on Circuits and Systems for Video Technology,
and Multimedia Tools and Applications (Springer Nature), 2024–2025.
Associate Member, Sigma Xi, The Scientific Research Honor Society (elected 2025).
|
|
Graduate Research Assistant (Aug 2024–May 2025) and Research Associate (Jan–May 2024)
Department of Computer Engineering, Rochester Institute of Technology — supervisor Prof. Dongfang Liu
|
|
Student Assistant (Jan–Dec 2023)
Department of Computer Science, Stony Brook University
|
|