[ SECTION / AI ]
LLMs
Frontier LLM technology, inference optimization, and training.
31 POSTS01
AI / 2026.08.12 / 21 MIN
Companies Offer 30K for Agent Engineers, Demo Writers Can't Get Hired
Companies can't hire production-ready agent engineers at 30K a month while demo writers can't find jobs. We break the gap into four skill blocks, each backed by real …
↗
02
AI / 2026.07.22 / 7 MIN
llama.cpp, vLLM, SGLang: Three Inference Engines, Three Completely Different Paths
llama.cpp, vLLM, and SGLang aren't competitors in the same lane — they're built on entirely different design assumptions. Which one fits your workload? The answer depends …
↗
03
AI / 2026.07.21 / 11 MIN
BigQuery AI.AGG: When GROUP BY Starts Calling Gemini
Google wired Gemini into a BigQuery aggregate. Each group now returns one natural-language answer. We walk through SQL semantics, hidden batching, cost pitfalls, and four …
↗
04
AI / 2026.07.16 / 6 MIN
Seven Questions About Kimi K3: China's Quiet Fable 5 Contender
Moonshot AI teased K3 with a 36-second video, but the real story was already unfolding on Chatbot Arena — where an anonymous model codenamed Kivine has been building 3D …
↗
05
AI / 2026.04.26 / 6 MIN
Sub-Agents vs Agent Teams: The Architecture Decision That Can Wreck Your System
The real question is how to choose between Sub-Agents and Agent Teams. Most people immediately think of multi-agent systems when tasks get complex, but that's often the …
↗
06
AI / 2026.04.19 / 5 MIN
Did Claude Opus 4.7 Secretly Raise Prices? 497 Developers Reveal the Truth
497 anonymous developer submissions reveal Claude Opus 4.7 consumes 37.3% more tokens than 4.6 on average, with API costs rising proportionally. Here's what caused the …
↗
07
AI / 2026.03.26 / 6 MIN
Your AI Agent Can Think, But It Can't Remember
AI agents can reason, plan, and converse—but forget everything once the session ends. The Ghost project solves this with a pure PostgreSQL-based infrastructure, turning …
↗
08
AI / 2026.03.24 / 5 MIN
Cramming a 400B Model into 48GB: The Magic Behind LLM in a Flash
An Apple paper from 2023 made it possible to run a 400 billion parameter model on an ordinary MacBook. The core technologies—MoE and quantization—hide an engineering …
↗
09
AI / 2026.03.23 / 6 MIN
90 Seconds of Waiting, Gone: How oMLX Buries Ollama on Mac
oMLX is built for Apple Silicon, using the MLX framework, SSD-backed KV cache, and continuous batching to cut TTFT from 90 seconds to 1-3 seconds in long-context …
↗
10
AI / 2026.03.19 / 17 MIN
Don't Build a Thousand Agents: How Ramp Automates Finance with One Agent
Ramp, America's fastest-growing enterprise finance platform valued at $32B with 50,000+ customers and $100B+ in annual transaction volume, chose a 'one Agent + a thousand …
↗
AI / 2026.08.12 / 21 MIN
Companies Offer 30K for Agent Engineers, Demo Writers Can't Get Hired
Companies can't hire production-ready agent engineers at 30K a month while demo writers can't find jobs. We break the gap into four skill blocks, each backed by real …
↗
02
AI / 2026.07.22 / 7 MIN
llama.cpp, vLLM, SGLang: Three Inference Engines, Three Completely Different Paths
llama.cpp, vLLM, and SGLang aren't competitors in the same lane — they're built on entirely different design assumptions. Which one fits your workload? The answer depends …
↗
03
AI / 2026.07.21 / 11 MIN
BigQuery AI.AGG: When GROUP BY Starts Calling Gemini
Google wired Gemini into a BigQuery aggregate. Each group now returns one natural-language answer. We walk through SQL semantics, hidden batching, cost pitfalls, and four …
↗
04
AI / 2026.07.16 / 6 MIN
Seven Questions About Kimi K3: China's Quiet Fable 5 Contender
Moonshot AI teased K3 with a 36-second video, but the real story was already unfolding on Chatbot Arena — where an anonymous model codenamed Kivine has been building 3D …
↗
05
AI / 2026.04.26 / 6 MIN
Sub-Agents vs Agent Teams: The Architecture Decision That Can Wreck Your System
The real question is how to choose between Sub-Agents and Agent Teams. Most people immediately think of multi-agent systems when tasks get complex, but that's often the …
↗
06
AI / 2026.04.19 / 5 MIN
Did Claude Opus 4.7 Secretly Raise Prices? 497 Developers Reveal the Truth
497 anonymous developer submissions reveal Claude Opus 4.7 consumes 37.3% more tokens than 4.6 on average, with API costs rising proportionally. Here's what caused the …
↗
07
AI / 2026.03.26 / 6 MIN
Your AI Agent Can Think, But It Can't Remember
AI agents can reason, plan, and converse—but forget everything once the session ends. The Ghost project solves this with a pure PostgreSQL-based infrastructure, turning …
↗
08
AI / 2026.03.24 / 5 MIN
Cramming a 400B Model into 48GB: The Magic Behind LLM in a Flash
An Apple paper from 2023 made it possible to run a 400 billion parameter model on an ordinary MacBook. The core technologies—MoE and quantization—hide an engineering …
↗
09
AI / 2026.03.19 / 17 MIN
Don't Build a Thousand Agents: How Ramp Automates Finance with One Agent
Ramp, America's fastest-growing enterprise finance platform valued at $32B with 50,000+ customers and $100B+ in annual transaction volume, chose a 'one Agent + a thousand …
↗