Writing on AI/ML research, engineering lessons, and ideas worth thinking through.
Notes from building an autonomous research agent, and the multiple ways a loop like this quietly runs away from you. Most of them are cheaper to design around than to discover at 2am.
Putting an LLM inside your software opens a new attack surface: prompt injection, leaked training data, poisoned model weights, and agents with more permissions than they should have. A walkthrough of the four risks that matter most, and the concrete mitigations for each.
Evaluating an LLM-based agent isn't the same as evaluating the text it produces, you have to judge its decision making, its tool use, and the path it took to get there. This post works through final-report, stepwise, and trajectory-based evaluation, then applies all three to a four-agent research pipeline.
Not every query needs your most expensive model. A look at the gateway layer that sits between your application and your providers, and at RouteLLM's approach — reduce the choice to strong or weak, output a win probability, and turn the threshold into a cost dial you can sweep.
Natural language is redundant, and LLMLingua exploits that: 20x prompt compression on GSM8K for 1.5 points of accuracy. This post covers perplexity-based token ranking, why instructions, questions, and demonstrations can't be squeezed equally, and the budget controller math that decides who gets cut.
Decoding is memory bound, not compute bound: on an A100, Llama 7B spends 7ms moving weights and 0.045ms actually computing, leaving the GPU 99.4% idle. This post traces where the time goes and covers batching, chunked prefill, PagedAttention, FlashAttention, and what's left when you're only calling an API.