Deep dives on projects and research: autonomous ground and underwater vehicles, perception and reinforcement learning.
6 posts
Dozens of LoRA runs from Qwen 1.7B to 27B, a 3,416-card gold corpus, and an eval stack that kept moving under me. The fine-tunes matched a cheap frontier API on quality, but serving them only pays off at roughly ten tables running 24/7, and a new Luna generation at half the price made the API the right thing to design around.
How Sage decides when to speak: three weeks of moving the FIRE/SILENT decision from an LLM to a distilled DistilBERT. The labels were wrong before the model was, a silence timer was hiding 91% of the transcript, and distilling the arbiter made precision free. Includes a CPU microbenchmark, the cost math, and the honest gap between offline 90/90 and live P0.84/R0.69.
Sixty days into building a live AI co-pilot for my D&D table: per-player streaming STT, a silence-gated two-call agent over MCP tools and hybrid RAG, an honest latency budget, the benchmarks that lied to me, and the production incident that made me collapse five code paths into one.
Leading the software for Paradigm's entry in the 30th IGVC: vision-based obstacle detection, mapping and ROS 2 Nav2 for an autonomous course with lane lines and obstacles.
My MSc research (AIIDE-2023): a shared, world-centric view built from many agents' cameras that helps RL agents learn in multi-agent games.
My engineering capstone: autopilot perception, navigation and control for a subsea resident AUV, trained in a Unity simulator with reinforcement learning.