Kimi K3 AI: Pricing, Benchmarks and Kimi Work
A clear guide to Kimi K3 pricing, architecture, benchmarks, Kimi Work, research uses, risks, and official access.
Kimi K3 AI: Moonshot's 1 Million Token Challenger
Pricing, architecture, benchmarks, Kimi Work, research uses, risks, and access

Quick answer
Kimi K3 is Moonshot AI's flagship model. It offers a 1 million token context window, strong coding and agent performance, and lower list prices than several premium rivals. It is promising, but its full weights and technical report were still pending on July 20, 2026.
Moonshot AI has launched Kimi K3, a large artificial intelligence model built for coding, research, and long tasks that need a lot of context. The model has a 1 million token context window, native visual input, and a Mixture of Experts design with 2.8 trillion total parameters. Moonshot says it can activate 16 of 896 experts for each token, which helps the system use a very large model without using every part at once. 1
Kimi K3 matters because it combines three features that buyers often want at the same time: strong performance, a very large context window, and lower API prices. Its official price is $3 per million uncached input tokens, $0.30 for cached input, and $15 per million output tokens. These prices are below the published rates for GPT 5.6 Sol, Claude Fable 5, and Claude Opus 4.8 in several common cases. 2456
The model is not an automatic winner in every task. Moonshot itself says that Kimi K3 still trails Claude Fable 5 and GPT 5.6 Sol in broad overall quality. At the same time, independent tests show that it can compete well in coding, web development, and knowledge work. The balanced view is clear: Kimi K3 is a serious frontier challenger, but teams should test it on their own work before making a final choice. 1789
What Is Kimi K3?
Kimi K3 is the current flagship model from Moonshot AI, a Beijing based artificial intelligence company. It was announced in July 2026 as a model for long horizon coding, deep reasoning, and complete knowledge work. The company has placed it across Kimi.com, Kimi Code, Kimi Work, and the official Kimi API. 123
These product names are related, but they are not the same thing. Kimi K3 is the foundation model. Kimi.com is the main web and mobile experience. Kimi Code is a coding agent for terminals and development tools. Kimi Work is a desktop agent that can use approved local files, run code, work in a browser, and create reports, documents, spreadsheets, slides, and websites. 23
Moonshot calls K3 an open 3 trillion class model. However, the full model weights were not yet available on Moonshot's public Hugging Face model list when this article was checked on July 20, 2026. Moonshot says the full weights will be released by July 27, 2026. Until that release and its licence are verified, the safest wording is that Kimi K3 is an announced open weight model with a public weights release still pending. 111
Inside the Kimi K3 Architecture
Kimi K3 uses a Mixture of Experts architecture. A simple way to understand this is to picture a large hospital with hundreds of specialists. A patient does not need every doctor at the same time. The system sends each case to the small group that is most useful. Kimi K3 works in a similar way. It has 896 expert parts, but it activates 16 experts for each token. 1
This sparse design can reduce the amount of active computing compared with a model that uses every parameter for every token. Still, the total model is extremely large. Moonshot recommends large accelerator systems for efficient deployment, so running the full model will not be practical for most individuals or small teams unless a hosting provider offers it as a service. 1
The architecture also uses Kimi Delta Attention, often called KDA, and Attention Residuals. KDA is designed to improve long context efficiency and reduce memory pressure. Attention Residuals let deeper layers choose useful information from earlier layers instead of treating every earlier signal in the same fixed way. Moonshot says these changes, together with its Stable LatentMoE design and training methods, improve scaling efficiency by about 2.5 times compared with Kimi K2. This is a company claim and should be read as such until the full technical report and independent replications are available. 112
Kimi K3 is also natively multimodal. Official API examples show text, image, and video input. That gives developers more room to build systems that study documents, charts, interfaces, screenshots, and visual evidence inside the same workflow. 2

Original concept animation for sparse expert routing. It is not a copy of a vendor diagram.
What a 1 Million Token Context Window Really Means
A context window is the amount of information a model can consider during one request or session. Kimi K3 supports about 1 million tokens. In simple terms, that can cover a very large set of documents, a long software project, or an extended agent history. 12
For a software team, this could mean reading many files from a large codebase before suggesting a change. For a researcher, it could mean comparing a collection of papers, notes, tables, and code in one task. For legal and finance teams, it could support work across large contracts, filings, policies, or company reports. A longer window can reduce the need to split material into many smaller requests.
However, context size is only a capacity limit. It does not promise perfect memory, perfect search, or correct reasoning across every token. Models can miss details in the middle of a long prompt, focus on the wrong evidence, or repeat errors from earlier steps. Important facts, calculations, citations, and decisions still need human checking.
Kimi K3 is also not the only model in the 1 million token class. GPT 5.6 Sol lists a 1,050,000 token window, while Claude Fable 5 and Claude Opus 4.8 list 1 million tokens. The difference is often more about price, quality, tool support, and how well each system uses long context than the headline number alone. 456
Kimi K3 Benchmarks and Real World Performance
Benchmark results help show where a model may be strong, but no single score can predict every real task. Moonshot reports strong gains in coding, agent work, knowledge tasks, and visual reasoning. It also states that K3 remains behind Claude Fable 5 and GPT 5.6 Sol in overall experience and broad capability. 1
Independent sources give a more useful picture. Artificial Analysis lists Kimi K3 with an Intelligence Index score of 57. Vals AI reported 74.7 percent for Kimi K3 on its industry focused index, just below Claude Fable 5 at 75.1 percent and above GPT 5.6 Sol at 73.1 percent when checked for this article. 78
Arena's WebDev leaderboard placed Kimi K3 first with a preliminary score of 1679, ahead of Claude Fable 5 and GPT 5.6 Sol in that specific front end development test. This result is important for people who build websites and interfaces, but it should not be treated as proof that K3 is best at all forms of coding. The result was marked preliminary, and leaderboard positions can change as more votes and models are added. 9
The right lesson is not that Kimi K3 defeats every rival. The stronger lesson is that it performs close enough to premium models to deserve a direct test. Teams should compare answer quality, task completion, latency, tool use, output length, and total cost on the exact work they plan to run.
Figure 2. Vals AI snapshot checked July 20, 2026. Benchmark scores can change and do not predict every task.
Kimi K3 vs GPT 5.6 Sol, Claude Fable 5, and Claude Opus 4.8
Kimi K3 competes most clearly on price and long context work. Its uncached input price is lower than GPT 5.6 Sol, Claude Fable 5, and Claude Opus 4.8. Its output price is also lower than all three. That can matter in agent tasks, because an agent may produce many tokens while planning, calling tools, checking results, and writing a final answer. 2456
Claude Fable 5 has the strongest broad positioning among the Anthropic models in this comparison. Anthropic describes it as its most capable widely released model for demanding reasoning and long horizon agent work. It costs $10 per million input tokens and $50 per million output tokens. 5
GPT 5.6 Sol has a slightly larger context window at 1,050,000 tokens and a broad tool set through OpenAI's API. Its standard rate is $5 per million input tokens, $0.50 for cached input, and $30 per million output tokens. OpenAI also states that requests above 272,000 input tokens are charged at higher rates for the full request, which can affect very large context jobs. 4
Claude Opus 4.8 sits between K3 and Fable 5 on standard price. It costs $5 per million input tokens and $25 per million output tokens. Anthropic says the full 1 million token window is billed at standard pricing. 6
A cheaper token price does not always mean a cheaper completed task. A model that needs more retries, more reasoning, or longer output may cost more in practice. The best comparison is cost per successful task, not cost per token alone.
Model comparison table
| Category | Kimi K3 | Claude Fable 5 | GPT 5.6 Sol | Claude Opus 4.8 |
|---|---|---|---|---|
| Provider | Moonshot AI | Anthropic | OpenAI | Anthropic |
| Context window | 1M tokens | 1M tokens | 1.05M tokens | 1M tokens |
| Input price | $3 per MTok | $10 per MTok | $5 per MTok | $5 per MTok |
| Cached input | $0.30 per MTok | Prompt cache rates vary | $0.50 per MTok | $0.50 read, $6.25 write |
| Output price | $15 per MTok | $50 per MTok | $30 per MTok | $25 per MTok |
| Long context price | No separate tier stated | Standard price across 1M | Higher above 272K input | Standard price across 1M |
| Main strength | Cost, context, coding | Broad top tier reasoning | Tools and broad platform | Agent work and enterprise use |
| Key caution | Weights and report pending | Highest standard price | Long context surcharge | Premium cost |
Figure 3. Standard API input and output prices, checked July 20, 2026. Cached rates and long context rules differ.
Kimi K3 Pricing
Kimi K3 uses three main token rates. Cached input costs $0.30 per million tokens. Uncached input costs $3 per million tokens. Output costs $15 per million tokens. Pricing was checked on July 20, 2026. 2
Cached input is useful when a large part of the prompt stays the same across many requests. A team may keep a product manual, research library, or codebase summary as a stable prompt prefix. Kimi's official documentation says context caching is automatic, so developers do not need to manage a separate cache ID for normal requests. 2
Pricing comparisons must separate cached input, uncached input, and output. Mixing these numbers can create a false result. Tool fees, extra searches, file processing, retries, and long outputs can also raise the final bill. Teams should test a small set of real tasks and record the total tokens and tool calls before estimating monthly cost.
What Is Kimi Work?
Kimi Work is Moonshot's desktop AI agent for knowledge workers. It is not the name of the Kimi K3 model itself. The app can work with folders that a user approves, run Python and shell tasks, automate browser actions through WebBridge, and schedule repeated jobs. 3
Moonshot also promotes Agent Swarm, which can coordinate up to 300 sub agents for complex jobs. In practice, access, limits, and concurrency can depend on the user's plan and the product mode. A large number of agents does not guarantee a better result. Good task design, clear permissions, and strong review steps still matter. 3
Kimi Work can create many kinds of output, including research reports, documents, spreadsheets, slide decks, charts, dashboards, and websites. This makes it useful for jobs that start with files and data but end with a polished deliverable.
The local connection also creates risk. A desktop agent may have access to private files, active browser sessions, or commands that can change data. Users should approve only the folders and accounts needed for the task. Sensitive work should follow company policy, data protection rules, and security review. No team should assume that all processing stays on the device unless the official privacy terms confirm that point for the exact feature being used.
How Researchers Could Use Kimi K3
Researchers may find Kimi K3 useful because it can combine long documents, code, data, and visual output in one workflow. A research team could load a set of papers, ask the model to map major themes, identify disagreements, and create a reading plan. It could then help clean data, write analysis code, and produce draft charts.
The model may also help with citation organization, research outlines, method checklists, and plain language summaries. Kimi Work can connect these steps by reading approved local files, running Python, and producing reports or presentations. 3
Human review remains essential. AI systems can invent references, misread tables, use the wrong statistical test, or write a confident claim that the data does not support. Researchers should verify every citation, calculation, dataset, code result, and interpretation. They should also follow journal, university, and funder rules for AI disclosure and authorship.
Kimi K3 should support scientific work, not replace scientific judgment. A good workflow gives the model clear source material, asks it to show evidence, tests its code, and requires a human decision before publication.
Limitations, Safety, and Legal Risks
Moonshot lists several limits for Kimi K3. The model can be sensitive to missing reasoning history, which may reduce quality when an agent system fails to pass the full conversation back to the model. The company also says the model can act too proactively when instructions are unclear. 1
Large context prompts can create their own problems. A long prompt may contain conflicting facts, hidden malicious instructions, old information, or private data. More context can also increase latency and cost. Teams should filter input, set clear task boundaries, and keep a record of the sources used.
Kimi Work adds operational risk because it can interact with files, code, browsers, and accounts. Use the least access needed. Keep backups. Require approval before deletion, purchase, publication, or account changes. Test agents in a safe workspace before allowing them to touch important systems.
There is also a licensing issue. Open weights and open source are not the same. Open weights means the trained parameters are available. Open source may also require code, training details, and a licence that allows broad use and modification. The exact K3 terms should be checked when the weights are released.
This article uses company and product names only for reporting and comparison. It does not suggest endorsement or partnership. Pricing, access, benchmarks, and licence terms can change, so readers should confirm current details before making a purchase or deployment decision.
How to Access Kimi K3
Users can access Kimi K3 through official Moonshot products. The main options are Kimi.com, Kimi Work, Kimi Code, and the Kimi API platform. Kimi Work is available for Windows and Apple silicon Mac computers according to current product documentation. 123
Developers can use the official API for text, image, video, tool use, and long context tasks. Researchers and companies that want to host the model themselves should wait for the official weights, technical report, and licence, then review hardware needs and security controls. 12
Avoid unofficial lookalike sites. Use links from Moonshot's official pages, and confirm the domain before entering payment details, uploading sensitive files, or creating an API key.
Is Kimi K3 a Genuine Frontier Challenger?
Yes, Kimi K3 is a genuine challenger, especially for coding, long context work, research support, and price sensitive agent tasks. Independent results show that it can compete near leading models, and Arena's current web development result is a strong signal for front end work. 789
It is still too early to call K3 the best model overall. The full weights and technical report were still pending when this article was updated. Some benchmark results come from Moonshot, and some independent rankings are preliminary. Quality can also change across languages, tools, prompts, and product surfaces.
The most practical conclusion is simple. Put Kimi K3 on the shortlist. Test it beside GPT 5.6 Sol, Claude Fable 5, and Claude Opus 4.8 with the same prompts and success rules. Measure completed work, not marketing claims. For many teams, K3's mix of context, capability, and price may be its strongest advantage.
Frequently Asked Questions
What is Kimi K3?Kimi K3 is Moonshot AI's flagship model for coding, reasoning, and knowledge work. It uses a large Mixture of Experts design, supports visual input, and offers a 1 million token context window.
Is Kimi K3 open source or open weight?Moonshot describes K3 as an open model, but its full weights were still scheduled for release by July 27, 2026 when this article was checked. The final licence and release package should be reviewed before calling it fully open source.
Does Kimi K3 support 1 million tokens?Yes. Official Kimi documentation lists a context window of about 1 million tokens. This is a capacity limit, not a promise of perfect recall or reasoning across every part of a long prompt.
Is Kimi K3 better than GPT 5.6 Sol?Kimi K3 is cheaper and performs strongly in several coding and agent tests. GPT 5.6 Sol remains stronger in some broad evaluations and has a mature tool platform. The best model depends on the exact task.
Is Kimi K3 cheaper than Claude Fable 5?At published API rates, yes. Kimi K3 lists $3 per million uncached input tokens and $15 per million output tokens. Claude Fable 5 lists $10 for input and $50 for output.
What is Kimi Work?Kimi Work is a desktop agent that can use approved local files, run code, automate browser tasks, schedule jobs, and create reports, documents, sheets, slides, and websites.
Can Kimi Work access files on a computer?Yes, it can work with local folders that the user approves. Users should limit access, protect sensitive data, keep backups, and require review before high impact actions.
Where can developers access the Kimi K3 API?Developers should use Moonshot's official Kimi API platform. The main official access routes are Kimi.com, Kimi Work, Kimi Code, and platform.kimi.ai.
Sources and Fact Check References
- Moonshot AI, Kimi K3 Tech Blog: Open Frontier Intelligence
- Kimi API Platform, Kimi K3 Quickstart and Pricing
- Kimi Help Center, Kimi Work Overview
- OpenAI, GPT 5.6 Sol Model Documentation
- Anthropic, Introducing Claude Fable 5
- Anthropic, Claude Models and Pricing
- Artificial Analysis, Kimi K3 Model Analysis
- Vals AI, Industry Benchmark Index
- Arena, WebDev Leaderboard
- Reuters, Moonshot Kimi K3 Launch Coverage
- Moonshot AI on Hugging Face, Current Model List
- Moonshot AI, Kimi Linear and Attention Residuals Repositories
Legal and Editorial Note
This is an independent editorial article. It is not sponsored by, affiliated with, or endorsed by the companies and publishers named above. Company and product names are used for identification and fair comparison. All prose, tables, charts, and animations on this page are original.


