Independent AI researcher: embeddings, evaluation, and attribution for deployed AI Author of the open-source GrandJury evaluation protocol (arXiv:2508.02926) Production ML in regulated finance, healthcare, and legal Master of Applied Data Science @ University of Michigan
Independent AI researcher working on evaluation and attribution for deployed AI — how we tell whether AI output is trustworthy, and who is accountable when it isn't. My work centers on embeddings, information retrieval, and human evaluation. At the heart of my work is a commitment to rational thinking and science-driven innovation — it's about asking the right questions, and making tech that truly helps people. When I'm not engaged in tech, I'm into music. Modern architecture and design also spark my curiosity.
Democratizing Evaluation - A Personal Statement
https://amiaconference.net/amia-2025-speaker-information/
HumanJudge | open, production human-evaluation infrastructure
https://github.com/memoirji-llc/stt-inference-evaluation-pipeline
AMIA 2025: Transcribing a Broken Record: Benchmarki...
Building Intuitions Towards Evaluation Metrics for NLP Tasks
GrandJury: A Collaborative Machine Learning Model Evaluation...
設計機器學習系統: 迭代開發生產環境就緒的ML程式 | 誠品線上
Comparing Attention Mechanisms in Transformers