Close Menu
    Facebook X (Twitter) Instagram
    Cloud Tech ReportCloud Tech Report
    • Home
    • Crypto News
      • Bitcoin
      • Ethereum
      • Altcoins
      • Blockchain
      • DeFi
    • AI News
    • Stock News
    • Learn
      • AI for Beginners
      • AI Tips
      • Make Money with AI
    • Reviews
    • Tools
      • Best AI Tools
      • Crypto Market Cap List
      • Stock Market Overview
      • Market Heatmap
    • Contact
    Cloud Tech ReportCloud Tech Report
    Home»AI News»Moonshot AI Releases Kimi K3: A 2.8 Trillion Parameter Open MoE Model With Kimi Delta Attention and 1M Context
    AI News

    Moonshot AI Releases Kimi K3: A 2.8 Trillion Parameter Open MoE Model With Kimi Delta Attention and 1M Context

    July 17, 2026
    Facebook Twitter Pinterest Telegram LinkedIn Tumblr WhatsApp Email
    Moonshot AI Releases Kimi K3: A 2.8 Trillion Parameter Open MoE Model With Kimi Delta Attention and 1M Context
    Share
    Facebook Twitter LinkedIn Pinterest Telegram Email
    synthesia


    Moonshot AI just released Kimi K3. It is a 2.8-trillion-parameter model with native vision and a 1-million-token context window. Moonshot calls it the world’s first open 3T-class model.

    What is Kimi K3?

    Kimi K3 is a sparse Mixture-of-Experts (MoE) model built on two architectural updates. Those are Kimi Delta Attention (KDA) and Attention Residuals (AttnRes). Both change how information flows across sequence length and model depth. K3 targets long-horizon coding, knowledge work, and reasoning.

    Moonshot team states K3 is the first open model to reach 2.8 trillion parameters. For nine of the past twelve months, Kimi models set the upper bound of open-model sizes.

    Moonshot is also direct about where K3 sits. Overall performance still trails the most powerful proprietary models, Claude Fable 5 and GPT 5.6 Sol. Across Moonshot’s own evaluation suite, K3 consistently outperformed other tested models.

    quillbot
    https://www.kimi.com/blog/kimi-k3

    The Architecture Underneath

    Kimi Delta Attention (KDA) is a hybrid linear attention mechanism. Moonshot states it enables up to 6.3x faster decoding in million-token contexts.

    AttnRes works along the other axis, which is depth. It selectively retrieves representations across depth rather than accumulating them uniformly. Moonshot states AttnRes delivers roughly 25% higher training efficiency at under 2% additional cost.

    Sparsity is the third lever. K3 uses Stable LatentMoE, effectively activating 16 of 896 experts. At that sparsity, routing and optimization become first-order challenges. Quantile Balancing derives expert allocation directly from router-score quantiles. That eliminates heuristic updates and a sensitive balancing hyperparameter. Per-Head Muon extends Muon by optimizing attention heads independently. Sigmoid Tanh Unit (SiTU) and Gated MLA improve activation control and attention selectivity respectively.

    Refined training and data recipes accompany those structural changes. Together they yield roughly 2.5x better overall scaling efficiency than Kimi K2.

    Those choices carry into serving. K3 applies quantization-aware training from the SFT stage onward. It uses MXFP4 weights with MXFP8 activations for broad hardware compatibility. Moonshot team recommends supernode configurations with 64 or more accelerators. Because KDA poses new challenges for prefix caching, Moonshot contributed an implementation to vLLM.

    Performance

    With the mechanics established, the published scores are easier to read. All K3 results use reasoning effort set to max. Harnesses differ per benchmark: KimiCode, Claude Code, or Codex.

    BenchmarkKimi K3Fable 5 (w/ fallback)GPT 5.6 SolOpus 4.8GLM-5.2DeepSWE67.570.073.059.046.2Program Bench77.876.877.671.963.7Terminal Bench 2.188.384.688.884.682.7FrontierSWE81.286.671.366.767.3SWE Marathon42.035.039.040.013.0BrowseComp91.288.090.484.3—Automation Bench30.829.129.727.212.9GPQA-Diamond93.592.694.191.091.2HLE-Full43.553.344.549.8—MMMU-Pro81.681.283.078.9—OmniDocBench91.189.885.887.9—

    Two caveats shape this table. ‘With fallback’ means requests Fable 5 refuses under its usage policy route to Opus 4.8. Also, BrowseComp used context compaction triggered at 300K tokens. Without that context management, K3 scores 90.4.

    So K3 leads Program Bench, SWE Marathon, BrowseComp, Automation Bench, and OmniDocBench. It trails Fable 5 on FrontierSWE and HLE-Full, and GPT 5.6 Sol on DeepSWE.

    Use Cases and Examples

    Use caseReported exampleRelies onRepo-scale engineeringLong sessions, minimal human oversightKimi Code, /modelVision in the loopIterating between code and live screenshotsVision, ms://<file-id>Research reproductionI–Love–Q relations: 20+ papers, 3,000+ lines of Python1M context, auto cachingDeep research reports42-year ASIC study: 2.8k+ fetches, 11k+ pagesKimi Work, WidgetsDocument parsingOmniDocBench score of 91.1Vision, structured output

    Moonshot team states one native multimodal architecture handles text, images, and video together.

    Access and a Minimal Call

    K3 is live on Kimi.com, Kimi Work, Kimi Code, and the API. Access runs through the OpenAI SDK against a Moonshot base URL.

    from openai import OpenAI
    import os

    client = OpenAI(api_key=os.environ[“MOONSHOT_API_KEY”],
    base_url=”https://api.moonshot.ai/v1″)

    completion = client.chat.completions.create(
    model=”kimi-k3″,
    reasoning_effort=”max”,
    messages=[{“role”: “user”, “content”: “Introduce Kimi K3 in one sentence.”}],
    )
    print(completion.choices[0].message.content)

    Four rules matter. reasoning_effort supports only max, and the K2.x thinking parameter must not be used. temperature, top_p, and n are fixed, so omit them. max_completion_tokens defaults to 131072 and reaches 1048576. In multi-turn and tool calls, return the complete assistant message.

    Pricing is flat, with no tiering by context length. Cache-hit input is $0.30/MTok, cache-miss is $3.00/MTok, and output is $15.00/MTok. The cache-hit rate is therefore the number to watch. Moonshot team reports above 90% cache hits in coding workloads.

    Key Takeaways

    • Kimi K3 is a 2.8T-parameter open MoE model activating 16 of 896 experts.
    • KDA, AttnRes, sparsity, and refined recipes give ~2.5x better scaling than K2.
    • K3 leads BrowseComp, SWE Marathon, OmniDocBench; trails Fable 5 on FrontierSWE and HLE-Full.
    • OpenAI-SDK compatible at $0.30/$3.00/$15.00 per MTok, with 1M context.

    Check out the Technical details and Try here . Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.

    Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us

    Asif Razzaq is the CEO of Marktechpost Media Inc.. As a visionary entrepreneur and engineer, Asif is committed to harnessing the potential of Artificial Intelligence for social good. His most recent endeavor is the launch of an Artificial Intelligence Media Platform, Marktechpost, which stands out for its in-depth coverage of machine learning and deep learning news that is both technically sound and easily understandable by a wide audience. The platform boasts of over 2 million monthly views, illustrating its popularity among audiences.



    Source link

    synthesia
    Share. Facebook Twitter Pinterest LinkedIn Tumblr Email

    Related Posts

    Meta, Microsoft, Nvidia, IBM, and others back open-weight AI

    July 26, 2026

    Meet the New Claude Opus 5: Frontier-Class Agentic Coding and Computer Use at Unchanged Opus Pricing

    July 25, 2026

    MIT projects selected for funding under US Department of Energy’s Genesis Mission | MIT News

    July 24, 2026

    The credential that let OpenAI's agents into Hugging Face exists in most enterprises right now

    July 23, 2026

    Google’s Gemini 3.6 Flash targets enterprise agent token costs

    July 22, 2026

    Alibaba’s Tongyi Lab Releases Qwen-Audio-3.0-TTS, a Hosted Text-to-Speech Model in Flash and Plus Tiers Across 16 Languages

    July 21, 2026
    bybit
    Latest Posts

    U.S. Spot Bitcoin, Ethereum ETFs Post Net Outflows as Institutional Demand Cools

    July 26, 2026

    2 Dividend Stocks to Hold Comfortably for the Next 5 Years

    July 26, 2026

    Meta, Microsoft, Nvidia, IBM, and others back open-weight AI

    July 26, 2026

    OpenAI Security Incident explained..

    July 25, 2026

    Dango Blockchain to Shut Down, Halt Perp DEX Trading

    July 25, 2026
    aistudios
    LEGAL INFORMATION
    • Privacy Policy
    • Terms Of Service
    • Social Media Disclaimer
    • DMCA Compliance
    • Anti-Spam Policy
    Top Insights

    The basics of AI image prompting

    July 26, 2026

    The Complete Guide to Making Cinematic AI Videos (2026)

    July 26, 2026
    coinbase
    Facebook X (Twitter) Instagram Pinterest
    © 2026 CloudTechReport.com - All rights reserved.

    Type above and press Enter to search. Press Esc to cancel.

    bitcoin
    Bitcoin (BTC) $ 65,455.00
    ethereum
    Ethereum (ETH) $ 1,960.29
    tether
    Tether (USDT) $ 0.999265
    bnb
    BNB (BNB) $ 575.45
    usd-coin
    USDC (USDC) $ 0.999703
    xrp
    XRP (XRP) $ 1.11
    solana
    Solana (SOL) $ 76.89
    tron
    TRON (TRX) $ 0.331934
    figure-heloc
    Figure Heloc (FIGR_HELOC) $ 1.03
    staked-ether
    Lido Staked Ether (STETH) $ 2,265.05