Muse Glimmer 30B — open weights that run on one GPU
Apache 2.0 release hints at Zuckerberg’s “personal intelligence” vision; paired with a public pledge to resume open-source model releases.
EN · TechCrunch ↗EN · Bloomberg ↗Meta’s Llama 4 open models remain the cheapest way to read big documents at scale (Scout’s 10M-token context from US$0.08/M), while the new proprietary Muse Spark line marks Meta’s frontier pivot.
Meta open-sourced the Llama family that much of the industry fine-tunes on. Its frontier work has moved to the proprietary Muse Spark line, with Llama 4 (April 2025) still the latest open Llama generation.
| Model | What it is for | Vendor list price |
|---|---|---|
| Llama 4 Scoutopen · 109B MoE (17B active, 16 experts) | 10M-token context for long-document RAG and synthesis — cheapest big-context option | from US$0.08 / US$0.30 per 1M (≈ RM0.33 / RM1.22, DeepInfra) |
| Llama 4 Maverickopen · 400B MoE (17B active, 128 experts) | Strongest open-weight generalist with vision, 1M ctx | US$0.17 / US$0.60 per 1M (≈ RM0.69 / RM2.45, DeepInfra) |
| Muse Spark 1.2 / Muse Codeproprietary frontier · 5 Aug 2026 | Meta’s closed frontier line plus its coding agent | See Meta API pages |
Last updated — Llama is open-weight: prices shown are the cheapest hosted providers (DeepInfra/Together), not a Meta list price. Approx RM at 1 USD ≈ RM4.08 (spot, 19 Sep 2026). Behemoth was never released — not listed. Figures are vendor public list prices in USD per 1M tokens, standard tier — no batch, cache, promo or peak-hour discount applied. Vendors change list prices without notice: confirm on the vendor’s own pricing page before budgeting.
Discontinued or fully superseded models, kept for historical reference. Active models — even previous-generation ones still on sale — always stay in the lineup above.
| Model | Era & what it was for | Historic list price |
|---|---|---|
| Llama 3.1 405BJul 2024 · 128K ctx | The largest classic Llama — widely fine-tuned through 2025 | ~US$3.50 / US$3.50 per 1M (≈ RM14.28) — historic hosted price |
| Llama 3.3 70BDec 2024 | The efficient 70B that many products still run | ~US$0.88 / US$0.88 per 1M (≈ RM3.59) — historic hosted price |
Recent announcements with outlet links (EN first, CN where available).
Apache 2.0 release hints at Zuckerberg’s “personal intelligence” vision; paired with a public pledge to resume open-source model releases.
EN · TechCrunch ↗EN · Bloomberg ↗Meta’s proprietary frontier line updates, alongside the coding-agent product.
EN · BenchLM provider page ↗One partner for every digital layer – explore what else we build, host and grow for Malaysian businesses.