
Evals show that a RAG system failed. Traces show whether retrieval, context construction, generation, or the eval itself caused it.
Most of what I enjoy is working around models, especially on systems that can gather knowledge and produce more accurate answers. That has pulled me into RAG, agentic search, fine-tuning, and agent harnesses, not as a checklist, but as different ways to give models new capabilities and make them useful in specific domains.
Before this, I spent a couple of years across [Public Goods], [DeFi], and fintech, working with Neverland, Giveth, Myosin, and others.
I’m also drawn to where software meets art and human behavior.
Most things we take for granted started with someone being curious enough to look closer.

Evals show that a RAG system failed. Traces show whether retrieval, context construction, generation, or the eval itself caused it.

Open-weights models, proprietary models, parameters, distillation, quantization, inference, deployment, and cost without the usual fog.

A compact map of NLP, NLU, NLG, tokenization, entity extraction, and sentiment before embeddings enter the picture.

How text representation moved from one-hot vectors, Bag of Words, and TF-IDF toward dense embeddings that capture meaning-like similarity.