I build production LLM, retrieval, and speech systems — and the evaluation infrastructure that tells you whether they actually work.
I own the AI function at a media technology company serving 100+ enterprise newsrooms, taking greenfield projects through production and setting AI strategy. That means fine-tuning and serving open-weight models on cloud GPUs we run ourselves, building hybrid retrieval, and putting evaluation gates in front of anything that reaches a customer.
Model discipline is my common thread across eighteen years of research and industry work in computational chemistry, speech recognition, and LLMs: measure honestly, test with real data, and solve the real customer pain. Most of what goes wrong in applied ML is an evaluation or user experience problem.