[ SECTION / AI ]
LLMs
Frontier LLM technology, inference optimization, and training.
32 POSTS01
AI / 2025.12.17 / 4 MIN
How I Use Skills to Turn My AI Assistant into a Loyal Workmate
Step-by-step guide to using Skills so Codex and Claude shift from generalists into assistants that follow your playbook.
↗
02
AI / 2025.12.13 / 7 MIN
ai-bindgen: I Tried This 'Cursed' Rust Proc-Macro That Lets an LLM Write Your Code at Compile Time
ai-bindgen is an experimental Rust proc-macro that calls LLM APIs at compile time to generate code. This article explores this compile-time AI tool, discusses the …
↗
03
//
AI / 2025.12.04 / 5 MIN
Mistral 3 Official Release: The Latest Work from Europe’s AI Giant
While OpenAI, Google, and Anthropic battle it out across the Atlantic, Europe’s AI power is quietly rising. On December 2, Paris-based Mistral AI officially released its …
↗
04
//
AI / 2025.11.11 / 4 MIN
Bring Kimi K2 Thinking Home with 247GB RAM: Dynamic 1-bit GGUF Field Notes
Step-by-step guide to running Unsloth's Dynamic 1-bit GGUF build of the 1T-parameter Kimi K2 Thinking model on high-end PCs, covering install, download, inference, …
↗
05
AI / 2025.10.30 / 4 MIN
Tokencake: Multi-Agent KV Cache Scheduling That Cuts vLLM Latency by Half
Beihang/Peking/Alibaba introduce Tokencake, a KV-cache-centric serving framework for multi-agent apps. With time+space scheduling plus CPU buffering and progressive GPU …
↗
06
//
AI / 2025.10.04 / 5 MIN
Agno-Go: Building AI Agents in Go - What's it Like Being 16x Faster than Python?
Rewriting AI Agent framework in Go brings 16x performance boost, 180ns agent startup, and only 1.2KB memory footprint - this is the extreme experience Agno-Go delivers
↗
07
AI / 2025.09.28 / 5 MIN
Is the AI Bubble Real? A Wall Street Veteran's Cold Take
Amid the red-hot AI investment frenzy, a Wall Street veteran offers a different perspective: we might be experiencing a new tech bubble. From power bottlenecks to …
↗
08
AI / 2025.08.08 / 4 MIN
Stop Copy-Pasting: Cursor CLI Puts an AI Coding Agent Right in Your Terminal
Say goodbye to constant context switching between browser and terminal. Cursor CLI injects a powerful AI Agent into your command line for chat-style code generation, file …
↗
09
AI / 2025.07.29 / 5 MIN
Zhipu AI GLM-4.5 Release: New Benchmark Challenging GPT-4's AI Coding Capabilities
Zhipu AI releases GLM-4.5 series models with 355B parameters, surpassing GPT-4.1 in SWE-bench tests, ranking 3rd globally in AI benchmarks, supporting code generation, …
↗
10
//
AI / 2025.01.29 / 4 MIN
DeepSeek Drops a Bombshell: V3.2-Exp Sparse Attention Mechanism Debuts, API Prices Slashed in Half Again
DeepSeek-V3.2-Exp released with groundbreaking DSA sparse attention technology, 2-3x faster inference, 30-40% memory reduction, and API prices cut by over 50%
↗
AI / 2025.12.17 / 4 MIN
How I Use Skills to Turn My AI Assistant into a Loyal Workmate
Step-by-step guide to using Skills so Codex and Claude shift from generalists into assistants that follow your playbook.
↗
02
AI / 2025.12.13 / 7 MIN
ai-bindgen: I Tried This 'Cursed' Rust Proc-Macro That Lets an LLM Write Your Code at Compile Time
ai-bindgen is an experimental Rust proc-macro that calls LLM APIs at compile time to generate code. This article explores this compile-time AI tool, discusses the …
↗
03
//
AI / 2025.12.04 / 5 MIN
Mistral 3 Official Release: The Latest Work from Europe’s AI Giant
While OpenAI, Google, and Anthropic battle it out across the Atlantic, Europe’s AI power is quietly rising. On December 2, Paris-based Mistral AI officially released its …
↗
04
//
AI / 2025.11.11 / 4 MIN
Bring Kimi K2 Thinking Home with 247GB RAM: Dynamic 1-bit GGUF Field Notes
Step-by-step guide to running Unsloth's Dynamic 1-bit GGUF build of the 1T-parameter Kimi K2 Thinking model on high-end PCs, covering install, download, inference, …
↗
05
AI / 2025.10.30 / 4 MIN
Tokencake: Multi-Agent KV Cache Scheduling That Cuts vLLM Latency by Half
Beihang/Peking/Alibaba introduce Tokencake, a KV-cache-centric serving framework for multi-agent apps. With time+space scheduling plus CPU buffering and progressive GPU …
↗
06
//
AI / 2025.10.04 / 5 MIN
Agno-Go: Building AI Agents in Go - What's it Like Being 16x Faster than Python?
Rewriting AI Agent framework in Go brings 16x performance boost, 180ns agent startup, and only 1.2KB memory footprint - this is the extreme experience Agno-Go delivers
↗
07
AI / 2025.09.28 / 5 MIN
Is the AI Bubble Real? A Wall Street Veteran's Cold Take
Amid the red-hot AI investment frenzy, a Wall Street veteran offers a different perspective: we might be experiencing a new tech bubble. From power bottlenecks to …
↗
08
AI / 2025.08.08 / 4 MIN
Stop Copy-Pasting: Cursor CLI Puts an AI Coding Agent Right in Your Terminal
Say goodbye to constant context switching between browser and terminal. Cursor CLI injects a powerful AI Agent into your command line for chat-style code generation, file …
↗
09
AI / 2025.07.29 / 5 MIN
Zhipu AI GLM-4.5 Release: New Benchmark Challenging GPT-4's AI Coding Capabilities
Zhipu AI releases GLM-4.5 series models with 355B parameters, surpassing GPT-4.1 in SWE-bench tests, ranking 3rd globally in AI benchmarks, supporting code generation, …
↗
10
//
AI / 2025.01.29 / 4 MIN
DeepSeek Drops a Bombshell: V3.2-Exp Sparse Attention Mechanism Debuts, API Prices Slashed in Half Again
DeepSeek-V3.2-Exp released with groundbreaking DSA sparse attention technology, 2-3x faster inference, 30-40% memory reduction, and API prices cut by over 50%
↗