Most people meet Kimi AI the same way. Someone mentions it can swallow a million tokens at once, they paste in a long PDF, ask it to summarise the thing, and then shrug. The summary is fine. Nothing about their day changes.
The model isn’t the problem. The workflow is. Long context is a capability, not a prompt, and it only pays off if you sequence the work properly. What follows is the routine I run when I need Kimi to do something genuinely useful with a stack of documents, including the prompts, the failure points, and the numbers.
What you’re actually working with
Kimi is built by Moonshot AI. The consumer chat app accepts PDFs, Word files, spreadsheets, slides, and plain text, and its long-context reputation is deserved. You can hand it entire reports instead of paragraphs, which changes the kind of question worth asking. If you want the background on how it earned that reputation, there’s a solid write-up on the assistant that reads a million tokens in one sitting.
Two facts shape everything below. Context length is a budget rather than a promise, and Kimi performs far better when you give it a job with a required output shape instead of a vague request.
Step 1: Upload first, ask second
Almost everyone types a question and then attaches a file. Flip that order. Upload the documents, then send one instruction that makes Kimi describe what it received.
Prompt: “List every file you have been given, the page count of each, and the top-level sections of each. Do not summarise the content yet.”
Last month I dropped in a 214-page master services agreement plus two amendments. Kimi returned a section map showing that the amendments renumbered clauses 7 through 11, something I would certainly have missed reading linearly. That map took about fifteen seconds and saved an hour of cross-referencing.
Why this step matters
Kimi builds a working index of the material before you interrogate it. Skip that and the first few answers drift toward whatever text sits near the beginning, because the opening pages dominate its attention. Ask for the map and you anchor the whole session.
Step 2: Ask narrow questions with a required format
Weak prompt: “What are the termination rights?”
Strong prompt: “Find every clause mentioning termination for convenience. For each one give the clause number, the page, the notice period in days, and a verbatim quote of no more than 25 words. If a clause is ambiguous, say so instead of interpreting it.”
The format constraint does the heavy lifting. Clause number, page, notice period, quote. Kimi cannot hand you a foggy paragraph when you have asked for four columns.
Tell it to refuse
Add one line to every research prompt: “If the documents do not answer this, write NOT FOUND and list which sections you checked.” Without that instruction, models fill gaps. With it, you get an honest null result, which is often the single most valuable finding in a session. Discovering that a contract is silent on data portability is worth more than a confident guess.
Step 3: Verify numbers with a second pass
Text extraction from a long context is reliable. Arithmetic on top of that text is not. I have watched a model correctly locate a 3.5% escalation clause and then compute eighteen months of it as though the year had ten months in it. Even systems built specifically to handle hard mathematics make slips a human spots in seconds, and watching researchers check each proof by hand after OpenAI’s ‘Astra’ solved 10 long-standing math problems was a useful reminder that verification never becomes optional.
So let Kimi find and quote. Do the arithmetic yourself or paste the figures into a calculator. Ask it to show its working when you want to audit the reasoning, never to replace the audit.
A 20-minute contract comparison, end to end
Five vendor agreements, 60 to 90 pages each, one table at the end of it.
- Upload all five, then request a file inventory.
- Define the columns: vendor, governing law, liability cap, auto-renewal, termination notice, data ownership.
- Ask for the table with a page reference in every cell.
- Run it again, this time asking Kimi to mark any cell it is less than confident about with a question mark.
- Open the flagged clauses yourself. In my last run, nine cells out of thirty were flagged and six of those genuinely needed a human read.
Notice the shape of it. A first pass for coverage, a second for uncertainty. You finish with a table you can defend in a meeting and a short list of items to check before you sign anything.
Using Kimi on a codebase
Consumer chat has limits. A 40,000-line repository will not fit usefully even in a generous window, because attention thins out in the middle. Pick three to five files, roughly 1,500 lines total, and state the goal plainly.
Prompt: “Here are router.js, auth.js, and db.js. The bug is that sessions survive a password change. Identify the cause, then give me a unified diff. Do not refactor anything unrelated.”
Asking for a diff rather than a rewrite keeps you in charge of the merge. If you are running heavier jobs, Kimi’s open-weight K2 model can be self-hosted or called through an API, which is where the token economics start making sense for batch processing.
Three prompts worth stealing
- The audit prompt. “Read all attached documents. List every obligation that carries a deadline, sorted by date, with page references.”
- The contradiction prompt. “Find statements in these documents that conflict with one another. Quote both, cite pages, and say which is more recent.”
- The devil’s advocate prompt. “Argue against the recommendation in section 4, using only evidence from the attached files.”
Where Kimi gets shaky
Four failure modes show up repeatedly. Numbers. Citations, where it will occasionally invent a plausible page reference. Anything dated after its training cutoff, unless web search is switched on. And the middle of very long inputs, where details get quietly dropped. The competitive pressure behind these models is enormous, and the pace is part of why everyone is freaking out about OpenAI and Anthropic’s race for dominance. New versions land fast enough that any specific limitation may be stale within a month.
Behaviour can also shift from the top down. Labs change what their systems will answer, sometimes with little warning, which is exactly what happened when OpenAI moved to put the safety brakes on Astra. Treat any tool as a temporary arrangement and test it against your own documents rather than trusting a review, including this one.
One practical guardrail: keep a running note of every claim Kimi makes that you later verify as wrong. Mine is eleven lines long after three months. It has taught me more about where to aim the model than any guide could.
The habit that makes long context pay off
Tools like this reward a specific kind of discipline. Inventory, narrow questions, shaped outputs, verification. Do those four things and a 200-page document becomes a two-hour job rather than a two-day one.
A routine that has stuck for me: every Monday morning, dump the week’s incoming documents into a single Kimi session, ask for the inventory and the deadline list, flag every ambiguous item, then export the table to a doc. Twenty minutes of setup. The rest of the week you spend answering questions instead of hunting for them.
Start with one document you already know inside out. You will spot the errors immediately, which is precisely the point, and you will learn the model’s habits faster than any tutorial can teach them to you.

