<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>WASTE on 梦兽编程</title><link>https://rexai.top/en/tags/waste/</link><description>Recent content in WASTE on 梦兽编程</description><generator>Hugo -- 0.163.3</generator><language>en</language><copyright>梦兽编程</copyright><lastBuildDate>Sat, 01 Aug 2026 09:30:00 +0800</lastBuildDate><atom:link href="https://rexai.top/en/tags/waste/index.xml" rel="self" type="application/rss+xml"/><item><title>2.78 Trillion Parameters, 29GB RAM, Half a Token Per Second: Kimi K3 Actually Runs on a Laptop</title><link>https://rexai.top/en/posts/2026-08-01-waste-kimi-k3/</link><pubDate>Sat, 01 Aug 2026 09:30:00 +0800</pubDate><guid>https://rexai.top/en/posts/2026-08-01-waste-kimi-k3/</guid><description>WASTE, a pure-C inference engine, squeezes the 2.78-trillion-parameter Kimi K3 into a 64GB laptop: a 982GB container streams expert weights from disk on demand, measured at 0.5 tok/s. Half a token per second is absurdly slow — but it proves running trillion-scale models locally is an engineering problem, not a money problem.</description></item></channel></rss>