How can we build LLM agents whose behaviour is reproducible, verifiable, and robust?
Reliability across changing models, tools, environments, prompts, and execution conditions.
Mathematics × Artificial Intelligence
I am Bowen Wang, a mathematics-trained AI researcher working at the intersection of LLM agents, AI evaluation, and mathematical methods for machine learning. I build evaluation frameworks that make modern AI systems reproducible, decomposable, and scientifically comparable.
Standardised end-to-end in the SURE-EVAL onboarding workflow — one shared protocol for runtime and inference validation.
Annotated through an LLM-assisted, human-verified pipeline for privacy and trustworthiness research.
Reproducible baselines for locating the exact evidence spans that justify a compliance decision.
Research direction
My work asks how principled modelling, reproducible execution, and agentic systems can help us diagnose and improve modern multimodal AI — not score it, but explain it.
Reliability across changing models, tools, environments, prompts, and execution conditions.
Optimisation, partial orders, and error decomposition that reveal failures hidden by aggregate scores.
From overlapping speech to agent tool trajectories, I study partially ordered, asynchronous, and weakly aligned information across modalities.
“Reproducibility is not a final checkbox. It is the infrastructure that makes evaluation scientifically useful.”
Publications & current work
A leakage-free stacking ensemble that turns noisy wearable vibration signals into reliable running-speed estimates — with the evaluation protocol designed so no information crosses the train/test boundary.
DOI 10.1109/EECR69522.2026.11549312 ↗Standardised 18 public speech and audio models across six task families, with reproducible runtime and inference validation — the infrastructure layer this site’s evidence strip is drawn from.
Project repository ↗A time-constrained partial-order alignment that separates lexical recognition errors from speaker-attribution errors, so multi-speaker ASR systems can be diagnosed instead of only ranked.
Manuscript under reviewResearch experience
Mar 2026 — Sep 2026
Supervisor: Prof. Kai Yu, IEEE Fellow
Shanghai Jiao Tong University
Jun 2025 — Mar 2026
Supervisor: Prof. Yan Zhang
Institute of Information Engineering, Chinese Academy of Sciences
Selected projects
Reproducible AI evaluation
A standardized onboarding and evaluation workflow for heterogeneous speech, audio, and multimodal models.
Mathematical evaluation
A decomposed metric separating lexical recognition errors from speaker-attribution errors in multi-speaker ASR.
Interpretable computer vision
Random Forest and ResNet18 pipelines with Grad-CAM and targeted occlusion studies for interpretable facial-region analysis.
Education
University of Bristol · Bristol, UK
University of Bristol · Upper Second-Class Honours
Let’s connect
If your group works on AI evaluation, agent reliability, mathematical machine learning, or multimodal systems, I would be glad to hear from you.
wwwency2003@outlook.com