A support team I worked with had 4,200 help articles and a search box nobody used. Their first instinct was a fine-tune. Their second — the one that actually shipped — was a retrieval pipeline about 180 lines long, built with LangChain over a weekend. It answered roughly 70% of internal questions correctly on day one, and the engineers could read every line of it.
That build is what this walkthrough covers. Not the theory, not the architecture diagram, but the seven concrete steps from a folder of PDFs to a working Q&A bot you can hand to someone else. Each step includes the decision that trips people up.
Step 1: Pin your versions before you write a line of code
LangChain reorganised its Python packages at the 0.2 release, then again in 0.3. Imports moved from one big library into langchain-core, langchain-community, and per-provider packages like langchain-openai. If you copy a tutorial that starts with from langchain.chat_models import ChatOpenAI, you’re reading 2023 code and it will throw an ImportError.
Set up a virtual environment and install the split packages directly:
pip install langchain langchain-openai langchain-community langchain-text-splitters faiss-cpu python-dotenv
Then freeze. Write the exact versions into requirements.txt the moment things work. LangChain moves fast enough that an unpinned rebuild three weeks later can behave differently, and debugging that is a miserable afternoon. Keep your API keys in a .env file and load them with python-dotenv rather than pasting them into notebooks you’ll later share.
Step 2: Load and split your documents without shredding context
Use DirectoryLoader with a file-specific loader for each type — PyPDFLoader for PDFs, UnstructuredMarkdownLoader for docs, TextLoader for plain text. The trap here is chunking. The default RecursiveCharacterTextSplitter at 1000 characters with 200 overlap is a reasonable starting point, but it will happily cut a table in half or split a code sample from its explanation.
For the 4,200-article corpus, 800 characters with 120 overlap worked better. Smaller chunks retrieve more precisely; larger chunks give the model more to reason over. Test both.
- Attach metadata at split time. Every chunk should carry its source filename and page number. Without it you can’t cite anything, and users won’t trust an answer they can’t verify.
- Clean before you chunk. Headers, footers, and page numbers repeated on 200 pages become noise in every embedding.
- Split on structure first. Markdown headers and HTML section tags make better boundaries than character counts.
- Log your chunk count. 4,200 articles should not produce 400 chunks. If it does, your loader is silently failing.
If you want the wider picture of what the framework is doing at each of these stages, the breakdown in how LangChain works under the hood maps cleanly onto this pipeline.
Step 3: Build the retrieval layer and persist it
Embed your chunks, drop them into a vector store, and expose a retriever:
vectorstore = FAISS.from_documents(chunks, OpenAIEmbeddings(model=”text-embedding-3-small”)), then retriever = vectorstore.as_retriever(search_type=”mmr”, search_kwargs={“k”: 4, “fetch_k”: 20})
Two things worth doing early. First, use MMR (maximal marginal relevance) rather than plain similarity search — it returns four different passages instead of four near-duplicates of the same paragraph, which is a common cause of confidently wrong answers. Second, call save_local() once the index is built. Re-embedding thousands of chunks on every script run costs real money and adds minutes to every test cycle.
Similarity search over embeddings is often enough. Add a reranker only when you’ve measured that retrieval is the bottleneck.
Step 4: Wire the chain together with LCEL
LangChain Expression Language is the part worth learning properly. It uses the pipe operator, the same way Unix does, and it’s composable in the same way:
chain = prompt | model | StrOutputParser()
For Q&A you need two inputs resolved in parallel — the user’s question and the retrieved context. RunnableParallel handles that. Because everything is a Runnable, you get .invoke(), .batch(), .stream(), and async variants for free, which matters a lot at step five.
One more upgrade: instead of parsing a string and hoping the model formatted it correctly, use model.with_structured_output(AnswerWithSources) and pass a Pydantic model. You get a validated object with the answer and a list of source chunks. If your team already leans on Pydantic everywhere, the approach described in type-safe AI agents built the FastAPI way is worth reading as a comparison — it swaps the chain for explicit Python and leans harder on validation.
Step 5: Add conversation history and streaming
A bot that forgets the previous sentence feels broken. Wrap the chain with RunnableWithMessageHistory and a session-based store keyed by user ID.
Use windowed memory, not an unbounded buffer. Keeping the last six exchanges is usually enough for support questions, and it caps your token spend. An unbounded history on a long session can quietly triple your cost per request by the twentieth turn.
Streaming is the other half of perceived quality. Swap .invoke() for .stream() and push tokens to the front end as they arrive. First-token latency drops from several seconds to a few hundred milliseconds, and users read the response as it forms. If you’d rather not assemble memory, retrieval, and tools by hand, Phidata’s batteries-included assistant pattern packages those same pieces with less wiring.
Step 6: Give the bot a tool — carefully
Retrieval answers “what does our policy say?” Tools answer “what’s the status of order 88213?” Decorate a normal Python function with @tool, write a docstring the model can actually understand, and bind it to the model with bind_tools(). The docstring is the interface — vague descriptions produce wrong tool calls far more often than bad prompts do.
Then set limits. Agents loop. Cap max_iterations at 5 or 6, add a timeout, and log every tool call with its arguments. If you find yourself writing a fifteenth integration for a SaaS API, a pre-built connector layer such as Composio’s tool integrations will save you the boilerplate.
Step 7: Test retrieval and generation as separate systems
This is the step everyone skips, and it’s the one that determines whether the bot survives contact with real users.
Write 25 question-answer pairs by hand. Then measure two things independently. Retrieval hit rate: did the correct chunk appear in the top 4? Answer accuracy: given good chunks, did the model produce the right response? A bot with 95% retrieval and 60% answer accuracy has a prompting problem. One with 50% retrieval has a chunking or embedding problem. Fixing the wrong one wastes days.
Wire up tracing so you can inspect the retrieved chunks behind any bad answer — LangSmith is the default choice, and running chains through it takes about four lines. When this prototype eventually needs to sit inside a governed data platform, the reliability patterns in building dependable AI systems on Databricks Mosaic AI are a good next read.
Knowing when to stop reaching for LangChain
LangChain earns its place when you’re prototyping, when you need a dozen integrations fast, and when LCEL’s streaming and batching do real work for you. It’s less compelling when your whole app is one prompt and one API call, or when you need strict compile-time guarantees about what the model returns.
There are now credible alternatives at every level. Managed agent platforms handle orchestration and hosting for you. Typed frameworks give you stronger contracts. Cloud-native options keep everything inside one vendor’s perimeter. Picking LangChain isn’t a permanent commitment — and the teams that get the most out of it treat the chain as one component in a system, not the system itself. The 180-line pipeline that answered 70% of those support questions did so because retrieval, chunking, and testing each got attention. The framework was the easy part.

