Mathematics × Artificial Intelligence

Measuring what machine intelligence actually does.

I am Bowen Wang, a mathematics-trained AI researcher working at the intersection of LLM agents, AI evaluation, and mathematical methods for machine learning. I build evaluation frameworks that make modern AI systems reproducible, decomposable, and scientifically comparable.

Incoming MSc Engineering with Artificial Intelligence · University of Bristol
BSc Mathematics, University of Bristol · Upper Second-Class Honours

Portrait of Bowen Wang wearing University of Bristol graduation dress
Bowen Wang
Bristol, UK
Open to PhD
opportunities · 2027

Public speech & audio models

Standardised end-to-end in the SURE-EVAL onboarding workflow — one shared protocol for runtime and inference validation.

Bilingual policy paragraphs

Annotated through an LLM-assisted, human-verified pipeline for privacy and trustworthiness research.

Micro-F1 · evidence localisation

Reproducible baselines for locating the exact evidence spans that justify a compliance decision.

Research direction

Evaluation as a structured scientific problem.

My work asks how principled modelling, reproducible execution, and agentic systems can help us diagnose and improve modern multimodal AI — not score it, but explain it.

LLM AgentsAI EvaluationMathematical Methods for AI Trustworthy AISpeech & Multimodal AI

How can we build LLM agents whose behaviour is reproducible, verifiable, and robust?

Reliability across changing models, tools, environments, prompts, and execution conditions.

Agent evaluationTool-use reliability

How can mathematical structure help us design better AI evaluation metrics?

Optimisation, partial orders, and error decomposition that reveal failures hidden by aggregate scores.

Partial-order DPMetric decomposition

How should we evaluate systems with complex temporal, structural, and multimodal dependencies?

From overlapping speech to agent tool trajectories, I study partially ordered, asynchronous, and weakly aligned information across modalities.

Temporal structureMultimodal evaluationLong-horizon interaction
“Reproducibility is not a final checkbox. It is the infrastructure that makes evaluation scientifically useful.”
— Working principle behind SURE-EVAL and SAER

Publications & current work

Research grounded in evidence, not headline scores.

  1. Accepted NCMMSC 2026

    SURE: Standardized Multimodal Model Onboarding and Reproducible Evaluation Framework

    Co-first author. Equal contribution

    Standardised 18 public speech and audio models across six task families, with reproducible runtime and inference validation — the infrastructure layer this site’s evidence strip is drawn from.

    Project repository ↗
  2. ICASSP 2027 · under review

    SAER: Decomposed Evaluation Metric for Speaker-Attributed ASR

    First author. Manuscript under review — not yet accepted

    A time-constrained partial-order alignment that separates lexical recognition errors from speaker-attribution errors, so multi-speaker ASR systems can be diagnosed instead of only ranked.

    Manuscript under review

Research experience

From trustworthy data to reproducible multimodal evaluation.

Mar 2026 — Sep 2026

Supervisor: Prof. Kai Yu, IEEE Fellow

Shanghai Jiao Tong University

LLM Agents · AI Evaluation · Speech & Multimodal AI

  • Co-developed a reproducible agentic evaluation framework spanning 18 public speech and audio models.
  • Designed structured evaluation for speaker-attributed ASR using temporal constraints and exact error decomposition.

Jun 2025 — Mar 2026

Supervisor: Prof. Yan Zhang

Institute of Information Engineering, Chinese Academy of Sciences

LLM-assisted AI · Trustworthy AI · Privacy & Security Analytics

  • Built a collaborative LLM annotation workflow covering 28,000+ bilingual privacy-policy paragraphs.
  • Developed reproducible baselines and evidence-span localisation reaching Micro-F1 = 0.91.

Selected projects

Systems, metrics, and interpretable machine learning.

Active

Reproducible AI evaluation

SURE-EVAL

A standardized onboarding and evaluation workflow for heterogeneous speech, audio, and multimodal models.

models passed one-shot onboarding after validation redesign
View integration notes ↗
Manuscript in review

Mathematical evaluation

SAER

A decomposed metric separating lexical recognition errors from speaker-attribution errors in multi-speaker ASR.

public SA-ASR pipelines evaluated on AISHELL-4 and AliMeeting
Research manuscript · ICASSP 2027 submission
Completed

Interpretable computer vision

Decoding Feline Pain

Random Forest and ResNet18 pipelines with Grad-CAM and targeted occlusion studies for interpretable facial-region analysis.

complementary landmark and full-image learning routes
View code and schema ↗

Education

A mathematical foundation for responsible AI.

2026 — 2027

MSc Engineering with Artificial Intelligence

University of Bristol · Bristol, UK

2023 — 2026

BSc Mathematics

University of Bristol · Upper Second-Class Honours

Let’s connect

I am interested in PhD opportunities starting in 2027.

If your group works on AI evaluation, agent reliability, mathematical machine learning, or multimodal systems, I would be glad to hear from you.

wwwency2003@outlook.com