Jul 18Meno's paradox, measured: Kimi K3's nervousness indexMinu ChoiJul 18Minu ChoiA Malayalam grammar puzzle only one document on earth contains, run through Opus 4.8 and Kimi K3 — and what their traces reveal about search, memory, and trust.
Jun 19The frontier is open-source todayHrishi OlickelJun 19Hrishi OlickelA detailed analysis and comparison of GLM-5.2 and Opus 4.8, both pushing forward state-of-the-art in timestamp-accurate diarization.
Jun 10Synthetic HiresHrishi OlickelJun 10Hrishi OlickelSynthetic Hires: High context agents in the critical path - the extended version of Hrishi's keynote at SuperAI 2026 Singapore.
May 17No Country for Old Code: Ship outcomes not toolsHrishi OlickelMay 17Hrishi OlickelExtended version, slides and interactive elements from Hrishi's talk at AI Engineer Singapore. Codons, sentinels, budgets, and the case for shipping outcomes instead of tools.
Mar 26Agentic Programming Patterns: 1Hrishi OlickelMar 26Hrishi OlickelGAN-inspired loops, automated review surfaces, and continuous improvement curves — patterns we've discovered building long-running agents with Hankweave.
Feb 26RISC won: Building towards Data AGIHrishi OlickelFeb 26Hrishi OlickelA deeply technical retrospective on the first year of Southbridge.
Feb 3Systems of Lasting ValueHrishi OlickelFeb 3Hrishi OlickelWhy does all code become legacy code, and what can we do about it in the age of AI?
Jan 15CCEPL-driven developmentHrishi OlickelJan 15Hrishi OlickelCoding agents are the new REPL: A place to test, but not to deploy. Here's the pattern of what we see that comes after.
Dec 31Antibrittle AgentsHrishi OlickelDec 31Hrishi OlickelThe bible on how to build reliable long-horizon AI agents that can work for hours or days without breaking. A guide to the architectural principles and practices that make agents antibrittle.
Dec 26Tap the BedMinu ChoiDec 26Minu ChoiWhat a Bambu 3D printer's bed-tapping ritual teaches us about building reliable AI agents - and why Hankweave forces agents to forget.
Dec 24We are so backHrishi OlickelDec 24Hrishi OlickelAgentic systems today feel like raw LLMs did two years ago - unpredictable, and sparking with the same AGI static. The agent (LLM + tools + loop) has become the new unit of computation, and it has dragged back every problem we thought we'd solved.
Dec 9Working with AI as a teamHrishi OlickelDec 9Hrishi OlickelFocusing on intent: how to keep AI from eroding the trust collaboration runs on.
Nov 30Hiring in 2025Hrishi OlickelNov 30Hrishi OlickelOpen-sourcing our take-homes and our hiring process, in the face of an increasingly agent-enabled world.
Jun 3Conducting smarter intelligences than me: new orchestrasHrishi OlickelJun 3Hrishi OlickelEverything we learned painstakingly orchestrating subagents made of every flagship model, manually reviewing their output, trying different methods of coordination, and figuring out what works.
May 6Everything I'll forget about EvalsHrishi OlickelMay 6Hrishi OlickelThe actually-actionable guide to AI evals that nobody else got around to writing, built from months of shipping real datasets and benchmarks. The puzzle at its center: how do you get a dumber intelligence to verify the output of a smarter one?
May 5Auto-generating Insanely Good ReportsHrishi OlickelMay 5Hrishi OlickelA repeatable report generation pipeline.
Dec 19Something Weird is Happening in AIHrishi OlickelDec 19Hrishi OlickelComparing models across small and big.
Dec 8Southbridge: Scalable data substratesHrishi OlickelDec 8Hrishi OlickelThe fundamental thesis underpinning Southbridge.
Oct 23Making use of the rest of our dataHrishi OlickelOct 23Hrishi OlickelSemistructured data has been the bane of my life - at my last company, half of all engineering hours went to wrestling it into shape. Why it is so stubborn, and how we started to fix it.
Jun 28A Comparative Analysis of Three AI Agent Approaches to a Complex Bug FixHrishi Olickel & Claude Opus 4Jun 28Hrishi Olickel & Claude Opus 4Three AI agents, one gnarly state-management bug in gemini-cli, three very different routes to a fix. Comparing Claude Code with and without subagents against Windsurf's Cascade, move by move.
Jun 15The Ghost in the Layout: A Debugging Story of Reactivity and RendersHrishi Olickel & Claude Opus 4Jun 15Hrishi Olickel & Claude Opus 4A React layout bug that vanished the moment you opened dev tools - the worst kind of ghost. A debugging story about reactivity, render timing, and chasing a non-deterministic gremlin back to its source.
Jun 9Mem0: Technical Analysis ReportHrishi Olickel & Claude Opus 4Jun 9Hrishi Olickel & Claude Opus 4An AI-generated teardown of Mem0, the memory layer for LLMs: its architecture, storage systems, and the novel pieces that make it work. Produced with the unsupervised report-writing workflow from 'new orchestras.'
Jun 3Claude Code: An analysisHrishi Olickel & Claude Opus 4 & Gemini 2.5 ProJun 3Hrishi Olickel & Claude Opus 4 & Gemini 2.5 ProA deep, AI-generated dissection of Claude Code's internals: dependencies, data structures, control flow, tools, and file editing. Generated by Opus 4 with help from most of the flagship models; the human write-up on how it was made lives alongside it.
Feb 19Latent space reasoning and the inner monologueHrishi Olickel & Gemini 2.0 Pro & OpenAI o1-proFeb 19Hrishi Olickel & Gemini 2.0 Pro & OpenAI o1-proWhat can the way LLMs reason in latent space teach us about our own inner speech? A wander through two papers on recurrent latent reasoning and the parallels to how minds talk to themselves.
Jan 10A review of editing with LLMsHrishi Olickel & Claude 3.5 SonnetJan 10Hrishi Olickel & Claude 3.5 SonnetA survey of how LLM systems actually edit code - Tabby, Claude, Aider, and the rest - and the one problem they are all really solving. How do you get a model to change exactly the right lines, reliably?
Jan 10Understanding Code Assistance in ZedHrishi Olickel & Claude 3.5 SonnetJan 10Hrishi Olickel & Claude 3.5 SonnetHow the Zed editor wires up AI code assistance under the hood - extraction, prompts, structured output, tool calling, and streaming. Written in Rust, explained with TypeScript analogies.
Dec 19Flash vs O1pro Long Context UnderstandingHrishi Olickel & Gemini Flash 2.0 & OpenAI o1-proDec 19Hrishi Olickel & Gemini Flash 2.0 & OpenAI o1-proGemini 2.0 Flash versus o1-pro on long-context comprehension, using Google's MMoE paper as the exam. The surprise is which model actually explains a complex architecture better.
Dec 9Better OCR with logprobsHrishi OlickelDec 9Hrishi OlickelCan a model's own confidence scores tell you where its OCR went wrong? Testing log-probabilities as an error detector for multimodal OCR on complex tables, with Gemini Flash and GPT-4o.
Dec 3Approaches for a priori estimation in holey breadHrishi Olickel & AI Cameos (Claude 3.5 Sonnet, GPT-4o, Gemini exp-1121, QwQ, OpenAI o1-preview)Dec 3Hrishi Olickel & AI Cameos (Claude 3.5 Sonnet, GPT-4o, Gemini exp-1121, QwQ, OpenAI o1-preview)Can a model estimate the calories and macros of a fancy holey bagel from a single photo? A bake-off across Claude Sonnet, GPT-4o, Gemini, and the reasoning models QwQ and o1-preview.
Nov 29Comparative Analysis of Three SongsHrishi Olickel & Gemini 1.5 ProNov 29Hrishi Olickel & Gemini 1.5 ProA multimodal model's close reading of three pieces from three traditions - a Carnatic kriti, an Urdu ghazal, and a Western folk-pop track - covering raga, structure, and cultural context. Each analyzed on its own, then against the others.
Nov 12Qwen2.5Coder-32b Code CritiqueHrishi Olickel & OpenAI o1-preview & Claude 3.5 Sonnet & GPT-4o & Qwen 2.5 CoderNov 12Hrishi Olickel & OpenAI o1-preview & Claude 3.5 Sonnet & GPT-4o & Qwen 2.5 CoderHow well can a 32B specialized coding model critique real code, next to o1-preview, GPT-4o, and Sonnet 3.5? The results say something surprising about specialized coding models.
Oct 14EntropixplainedHrishi Olickel & Claude 3.5 SonnetOct 14Hrishi Olickel & Claude 3.5 SonnetA guide to Entropix, _xjdr's context-aware sampler that tunes generation on entropy and attention rather than a fixed temperature. What it does, how the code works, and why it can improve coherence and cut hallucinations.
Sep 30OCR with GOT and SonnetHrishi Olickel & Minu ChoiSep 30Hrishi Olickel & Minu ChoiComplex Indian electoral tables that stumped Document Intelligence and Sonnet on their own, read at 98.79% accuracy. The trick: pair GOT's raw extraction with Claude Sonnet as a corrector.