Mohammad Dastgheib
  • Evaluation
  • Work
  • Research
  • Toolkit
  • Publications
  • CV
  • 35mm Film

Research Scientist at thebridge

Measuring what AI systems actually do

I design experiments, rubrics and benchmarks for evaluating models and human–AI decisions, grounded in psychophysics, rater reliability and Bayesian modeling.

View portfolio Download Résumé Download CV
Mohammad Dastgheib
01 / Work5Selected case studies
02 / Evidence2Published works + active manuscripts
03 / TrainingPhDCognitive neuroscience, UC Riverside
04 / Off-hours35mmAnalog photography

Industry experience

Research on how people and systems decide together

Current work on how humans and AI agents deliberate and decide together, with earlier experience supporting research operations behind surgical-safety technology.

thebridge logo

Research Scientist

thebridge Inc. · San Francisco, CA · Sep 2026 – Present

  • Design measurement frameworks for scoring the quality of group deliberation from conversation transcripts: construct definitions, behaviorally anchored rating scales, and explicit rules for when an item cannot be assessed.
  • Plan validation of human raters and LLM judges: reliability on ordinal ratings, human versus model agreement, and tests of whether an automated judge can stand in for a human annotator.
  • Treat preprocessing as a source of measurement error: how splitting a conversation into units changes scores, tested with sensitivity analyses rather than assumed away.
  • Build composite metrics with stated value choices and robustness checks across reasonable alternative specifications.
  • Write implementation specs and automated sanity checks so evaluation pipelines can be audited.
Surgical Safety Technologies logo

Research Assistant I

Surgical Safety Technologies · Toronto, ON · 2021

  • Supported instance-segmentation workflows for surgical tool detection and tracking.
  • Evaluated annotation quality across analyst teams using inter-rater reliability (κ and ICC).
  • Synthesized ML-for-surgical-safety research to support internal research operations.

Education

Training across neuroscience and psychology

A research foundation spanning perception, cognition, quantitative methods, and human behavior.

Neurotree profile (academic genealogy)

University of California, Riverside logo

PhD, Psychology & Cognitive Neuroscience

UC Riverside · Expected 2026

Queen's University logo

MSc, Psychology

Queen's University · 2020

York University logo

Hon. BSc, Biological Sciences & Psychology

York University · 2017

Selected work

Research that changes what a system does

Evaluation comes first. Each case study shows the question, method, tradeoff, and product implication, not just the final artifact.

Illustration of a measurement audit: bars, a magnifying glass, and calipers

01 · Model evaluation / grader audit

Audit the grader before trusting the effect

An apparent prompt-sensitivity gap was the regex, not the models.

Qwen 2.5 1.5B Instruct and Llama 3.2 1B Instruct wrote the answers. A rule-based regex graded them, not an LLM and not a person. The gap moved with the prompt wording until the audit showed the grader was scoring format, not arithmetic.

Grader auditExperiment designMeasurement
Architecture of a remote XR interaction study

02 · XR interaction / remote testbed

Why gaze produced more selection errors

99.2% of gaze errors were slips; hand input reached 5.15 bits/s.

Built a remote testbed, compared hand and simulated gaze with Fitts' law and NASA-TLX, then modeled failure dynamics with a Bayesian LBA. The gaze condition was simulated, not evaluated in a headset.

Experiment designBayesian LBAReact / TypeScript
Surgeon cognitive dashboard

03 · Human factors / proof of concept

Designing alerts around cognitive risk

Targeted ≤0.6 alerts/min with calibrated state probabilities.

Prototyped a training dashboard that turns pupil, HRV, and grip features into instructional thresholds. All validation used synthetic sessions; clinical effectiveness remains untested.

Human factorsCalibrationXGBoostR Shiny
Concept chart of dual-task performance

04 · Research roadmap / in progress

Locating the failure point in dual-task performance

Separating evidence quality, response caution, and motor strain.

A staged dissertation program connecting psychophysics, pupillometry, drift-diffusion modeling, and grip-force signals. Industry applications shown in the case study are concepts, not deployed systems.

PsychophysicsPupillometryDrift diffusion
Declarative memory study results

05 · Foundational research / NSERC

Testing whether brief meditation changes memory

60+ participants across behavioral and EEG measures.

Led a master's thesis from study design through analysis. The findings come from a controlled research sample and should not be read as a general product-effectiveness claim.

Mixed methodsEEGMemory

All case studies

Technical toolkit

Methods selected for the decision

I work across evaluation, measurement, inference, and implementation so the research question survives contact with the product.

01 / Evaluation

Test the claim

  • Behaviorally anchored rubrics
  • Human rater reliability (κ, ICC)
  • Grader and pipeline audits
  • Robustness across specifications
  • Human × human and human × agent studies
  • LLM-judge validation, in progress
02 / Quantitative

Model behavior

  • Psychophysics & signal detection
  • Bayesian DDM / LBA
  • Mixed-effects models
  • ML, calibration & SHAP
03 / Physiological

Measure latent state

  • Pupillometry & eye tracking
  • Grip-force & tremor analysis
  • EEG & psychophysiology
  • Artifact control & QC
04 / Implementation

Make it testable

  • React / TypeScript systems
  • R Shiny dashboards
  • Python, R, PyMC & Stan
  • Reproducible workflows

Start a conversation

Need evaluation results you can defend?

I am open to Research Scientist (Model Evaluation), AI Evaluation / Benchmarking, Human Evaluation Research, Quantitative UXR, and Human Factors roles. Authorized to work in the U.S. without sponsorship.

Email meDownload RésuméDownload CV
Back to top

© 2026 Mohammad Dastgheib  ·  AI Model Evaluation · Human Factors · Quantitative UX Research