Projects

Tools, prototypes, and system writeups around corpus audits, retrieval evaluation, agent traces, and practical LLM workflows.

  1. RAG Evaluation HarnessActive · 2026

    A repeatable loop for measuring retrieval quality, citation grounding, and refusal behavior in a small knowledge-base assistant — built so each run leaves behind debuggable evidence.

    TypeScriptAstroEmbeddingsVector SearchEval Sets