Research Scientist
- Design measurement frameworks for scoring the quality of group deliberation from conversation transcripts: construct definitions, behaviorally anchored rating scales, and explicit rules for when an item cannot be assessed.
- Plan validation of human raters and LLM judges: reliability on ordinal ratings, human versus model agreement, and tests of whether an automated judge can stand in for a human annotator.
- Treat preprocessing as a source of measurement error: how splitting a conversation into units changes scores, tested with sensitivity analyses rather than assumed away.
- Build composite metrics with stated value choices and robustness checks across reasonable alternative specifications.
- Write implementation specs and automated sanity checks so evaluation pipelines can be audited.




