Just Launched Gruve PulseAI Platform, your private AI infrastructure, production-ready in under 2 weeks.PulseAI is live — private AI, ready in 2 weeks.

See PulseAI
Blog

Kimi K3 hands-on: Cheap per token, expensive per task

August 11, 2026

Under the same lowest-tier usage limit, Kimi K3 could not finish a simple full-stack project in one quota cycle. Opus 4.8 and GPT-5.6 both could. The code K3 produced was not clearly worse. It just took more tokens and more time to get there.

That single result is the most useful thing I can tell you about Kimi K3, because it is the one thing the leaderboards do not show you.

Moonshot AI released K3 on July 16, with open weights following on July 27. It is a 2.8-trillion-parameter Mixture-of-Experts model with roughly 104 billion parameters active per token, a 1-million-token context window, and an always-on reasoning mode. It scores 57 on the Artificial Analysis Intelligence Index, top four overall and first among open-weight models. At $3.00 per million input tokens and $15.00 per million output tokens, the headline positioning is obvious: frontier-class intelligence at open-weight prices.

I spent time with it from Japan, on the kind of work our enterprise customers actually run. I want to be clear about scope: this is hands-on feedback, not a rigorous benchmark. Moonshot paused new consumer subscriptions on July 19, three days after launch, when demand pushed their GPUs to capacity. That capped how much I could run. Treat what follows as a practitioner’s read, not a measurement.

Where K3 is genuinely strong: understand and compress

I tested general office workloads: PowerPoint, Word, and Excel generation, summarization, information extraction, and Chinese, English, and Japanese writing and comprehension.

Output quality was quite competitive. For general generation tasks, accuracy was decent and the structure was mostly reasonable. Its strongest area was summarization and extraction. In several cases it captured almost all key information with very few omissions, the kind of result where you read the output, go back to the source, and cannot find the thing it dropped.

This is where I think the parameter count earns its keep. For “understand and compress” tasks, where the model has to hold a large input in view and decide what matters, scale helps. K3 performed very well there.

Multilingual: solid, with a Japanese accent problem

Chinese and English were both solid. Japanese was generally accurate, but occasionally the wording or sentence patterns felt slightly unnatural. It was not a comprehension issue, more like the output was understandable but not fully native.

That distinction matters for deployment. A model that misunderstands Japanese input is unusable. A model that understands correctly but produces slightly stiff Japanese is fine for internal summarization and extraction, and risky for customer-facing copy without a human in the loop.

Among the models I tested, GPT still gave me the best overall multilingual impression. Claude and K3 felt roughly comparable, with K3 possibly slightly ahead in a few cases.

Where it breaks down: multi-turn work

The weakness showed up in multi-turn coding, and it showed up hard.

Same task, same tier, same quota. K3 ran out before finishing. Speed and token consumption were the problem, not code quality. I used a medium reasoning-effort setting rather than the highest one, and it still felt slow, so I do not think this reduces to a misconfigured effort level.

My guess is that this relates to the thinking-history sensitivity Moonshot flags in their own materials. If part of the reasoning history is truncated during a multi-turn task, quality drops. Which means you need to preserve more reasoning context across turns. Which means that once a task becomes multi-step, token usage compounds fast.

I checked Chinese forums and social media afterward, and most of the feedback matched my impression. Quality assessments were subjective and varied. But “K3 is slow and token-intensive” was close to consensus.

Third-party measurement points the same direction. Artificial Analysis clocks K3 at roughly 33 output tokens per second and records it generating about twice the peer median in output tokens to complete the same index tasks. Because K3’s thinking pass is always on and every reasoning token bills as an output token at $15 per million, the token count is not an efficiency footnote. It is the invoice.

The cost story does not hold up the way the price card implies

Compared with DeepSeek, K3 does not feel especially cheap in practice. Given how fast it consumes tokens, the value-for-money story is weaker than the positioning suggests.

This is the gap worth internalizing: $3 per million input tokens is a unit price, not a cost. Your cost is unit price multiplied by tokens consumed to complete the task, and K3 moves the second number in the wrong direction. A model that costs a fifth as much per token and spends two to five times the tokens on reasoning is not delivering a fifth of the cost. On multi-turn agentic work, where reasoning context has to be preserved across turns and failed tool calls trigger retries, that multiplier compounds rather than averaging out.

Cached input at $0.30 per million helps if your workload re-sends the same long context repeatedly. It does nothing for output-heavy reasoning traces, which is exactly where K3 spends.

Is the open-versus-closed gap closing?

On output quality, yes, and faster than I expected. For summarization and information extraction specifically, K3 is competitive with the closed frontier models I use daily. If your workload is understand-and-compress, the open-weight option is now a real option.

On efficiency and cost, no. I do not think that gap has narrowed much at all. It is the clearest weakness I observed, and it is the one that determines whether a model survives contact with a production budget.

So the honest answer is that “open models have caught up” is true on one axis and not yet true on the other. Which axis you care about depends entirely on your workload shape.

What this means for enterprise deployment

For most enterprise voice and text processing: summarization, extraction, classification, routing, the interesting question is not whether K3 beats a closed frontier model on an index. It is whether you need a 2.8-trillion-parameter model at all. On a large share of that work, a smaller model finishes faster, costs less per completed task, and can run somewhere you control.

The pattern we keep seeing at Gruve is that model selection gets made on leaderboard position and per-token price, then the inference bill arrives shaped by tokens-per-task and turns-per-workflow, numbers nobody measured before signing. Open weights make this worse, not better, because now the deployment decision, the serving stack, and the capacity planning are yours too.

That is the work: benchmarking candidate models against your actual workloads rather than public indices, routing traffic so heavyweight reasoning is reserved for the tasks that need it, and running the inference layer so unit economics stay predictable as usage scales. Gruve does that across readiness, architecture, implementation, and managed operations.

If you are evaluating open-weight models for production, bring us your real workloads. We will help you find out what they actually cost before you commit. [Contact Gruve →]

Yutong Dai is based in Japan and works on enterprise AI deployment at Gruve.

Unlock your
true speed to scale

Accelerate what data and AI can do together.

Before you go - don’t miss what’s next in AI.

Stay ahead with Gruve’s monthly insights on trusted AI, enterprise data, and automation.