What I Mean by Evaluation
How I think about AI evaluation: defining quality on purpose, measuring held-out behavior with uncertainty, and failing closed when the evidence is not strong enough to trust.
How I think about AI evaluation: defining quality on purpose, measuring held-out behavior with uncertainty, and failing closed when the evidence is not strong enough to trust.
How I used an LLM-as-a-judge harness at Tresic as a production deploy gate across regions: why the gate had to sit in the decision path, how open-weight multi-region trust raised the bar, and why failing closed mattered more than a flattering score.
Why I leave room for curiosity, rabbit holes, and ideas that do not need to become a brand.
How learning about markets shaped my approach to evidence, risk, timing, and tradeoffs.
What I look for in anime, manga, light novels, visual novels, and psychological storytelling.
How I found Frenchcore and what remixing taught me about reinterpretation and making things from scratch.
What a decade of swimming taught me about consistency, nonlinear progress, and changing environments.
Why I enjoy finding bottlenecks, comparing tradeoffs, and building systems that survive their first use case.
Why I want room in my life for work, music, stories, exercise, and curiosity without turning everything into optimization.