梦兽编程
AI_SUITE

[ SECTION / AI ]

LLMs

Frontier LLM technology, inference optimization, and training.
31 POSTS
01 AI / 2026.08.12 / 21 MIN Companies Offer 30K for Agent Engineers, Demo Writers Can't Get Hired Companies can't hire production-ready agent engineers at 30K a month while demo writers can't find jobs. We break the gap into four skill blocks, each backed by real … 02 AI / 2026.07.22 / 7 MIN llama.cpp, vLLM, SGLang: Three Inference Engines, Three Completely Different Paths llama.cpp, vLLM, and SGLang aren't competitors in the same lane — they're built on entirely different design assumptions. Which one fits your workload? The answer depends … 03 AI / 2026.07.21 / 11 MIN BigQuery AI.AGG: When GROUP BY Starts Calling Gemini Google wired Gemini into a BigQuery aggregate. Each group now returns one natural-language answer. We walk through SQL semantics, hidden batching, cost pitfalls, and four … 04 AI / 2026.07.16 / 6 MIN Seven Questions About Kimi K3: China's Quiet Fable 5 Contender Moonshot AI teased K3 with a 36-second video, but the real story was already unfolding on Chatbot Arena — where an anonymous model codenamed Kivine has been building 3D … 05 AI / 2026.04.26 / 6 MIN Sub-Agents vs Agent Teams: The Architecture Decision That Can Wreck Your System The real question is how to choose between Sub-Agents and Agent Teams. Most people immediately think of multi-agent systems when tasks get complex, but that's often the … 06 AI / 2026.04.19 / 5 MIN Did Claude Opus 4.7 Secretly Raise Prices? 497 Developers Reveal the Truth 497 anonymous developer submissions reveal Claude Opus 4.7 consumes 37.3% more tokens than 4.6 on average, with API costs rising proportionally. Here's what caused the … 07 AI / 2026.03.26 / 6 MIN Your AI Agent Can Think, But It Can't Remember AI agents can reason, plan, and converse—but forget everything once the session ends. The Ghost project solves this with a pure PostgreSQL-based infrastructure, turning … 08 AI / 2026.03.24 / 5 MIN Cramming a 400B Model into 48GB: The Magic Behind LLM in a Flash An Apple paper from 2023 made it possible to run a 400 billion parameter model on an ordinary MacBook. The core technologies—MoE and quantization—hide an engineering … 09 AI / 2026.03.23 / 6 MIN 90 Seconds of Waiting, Gone: How oMLX Buries Ollama on Mac oMLX is built for Apple Silicon, using the MLX framework, SSD-backed KV cache, and continuous batching to cut TTFT from 90 seconds to 1-3 seconds in long-context … 10 AI / 2026.03.19 / 17 MIN Don't Build a Thousand Agents: How Ramp Automates Finance with One Agent Ramp, America's fastest-growing enterprise finance platform valued at $32B with 50,000+ customers and $100B+ in annual transaction volume, chose a 'one Agent + a thousand …