π I build software and ML systems β .NET and TypeScript web apps on one side, PyTorch and retrieval on the other
π― Iβm looking to collaborate on projects where the measurement matters as much as the model
π± Iβm currently learning retrieval evaluation, uncertainty calibration, and MCP tooling
π¬ Ask me about RAG architectures, calibrated prediction intervals, or making a free GPU do real work
β‘ Fun fact: I never miss a Grand Prix β and I made my model publish its picks before lights out, so it canβt armchair-quarterback either
I build systems that make their own internals visible. glassbox runs seven RAG architectures over one designed corpus and records every intermediate step, so you can see exactly where they diverge β the strongest reached 0.957 recall, the weakest 0.326. fortune-teller gives up on predicting stock prices and predicts uncertainty instead, beating its baseline on 15 of 15 tickers. Pitwall commits Formula 1 predictions to git before each race and scores them in public: it ties the "pole sitter wins" rule on accuracy and wins on calibration.
The rest is pinned below β .NET web apps, a parameter-efficient fine-tuning benchmark, and a handful of models with their evaluation written up in full.



