Most CrewAI tutorials stop at two agents that wave at each other and print “Hello, world.” That teaches you the API shape and nothing about the part that actually bites: getting three or four agents to hand work down a chain without dropping facts, repeating themselves, or burning through your API budget on the second task.
So let’s build something with a real deliverable. Target: a competitive research crew that takes a product name, digs up what three rivals are charging, analyses the gaps, and writes a one-page brief with sources. Roughly 60 lines of Python. Here’s the whole path.
Start from the deliverable, not the framework
Before you name a single agent, write down what lands on your desk at the end. For this crew it’s a markdown brief with a comparison table, three bullets on positioning gaps, and a source link for every claim.
That decision cascades. A table with required citations tells you the researcher needs search and scraping tools, the analyst needs to see the researcher’s raw output, and the writer needs a strict format spec. If you skip this step you end up with three agents that all sound like slightly different chatbots. CrewAI rewards role clarity more than any other single thing — the framework’s core bet is that agents with distinct roles and goals outperform one big prompt, and it only pays off if the roles are genuinely distinct.
Install and set up the project
CrewAI is plain Python. The CLI scaffolding is worth using once so you can see the folder layout, but for a crew this size a single file is easier to debug.
pip install crewai crewai-tools
crewai create crew competitor-brief # optional scaffold
You’ll need an LLM key in your environment. Set OPENAI_API_KEY (or point at Anthropic, or a local model through Ollama) and a search key if you’re using a Serper or Brave tool. Run a one-line smoke test before writing any agents — a bad key produces a confusing error halfway through a task otherwise.
Write agents with a narrow point of view
The role, goal, and backstory fields are not decoration. They’re injected into the system prompt and they decide how the agent judges its own work. Vague backstories produce vague output.
The researcher
from crewai import Agent
from crewai_tools import SerperDevTool, ScrapeWebsiteTool
search = SerperDevTool()
scrape = ScrapeWebsiteTool()
researcher = Agent(
role="Market Researcher",
goal="Collect verifiable pricing and positioning facts about rivals to {product}",
backstory=(
"You wrote analyst notes for eight years and you never cite a number "
"you cannot trace to a primary source. If a page is stale, you say so."
),
tools=[search, scrape],
max_iter=8,
verbose=True,
)
max_iter=8 is a guardrail, not a suggestion. Without it, a confused agent will happily loop through search calls until your bill gets interesting.
The analyst
Give this one no tools at all. That’s deliberate — an analyst that can search will start researching, which duplicates work and creates two competing fact sets.
The writer
Temperature down, format spec explicit. This agent’s only job is turning structured findings into prose a human will read on a Monday morning.
Turn the job into tasks with explicit outputs
This is where most crews quietly fail. expected_output is not a hint; it’s the contract the agent grades itself against and, in sequential mode, the payload passed downstream.
from crewai import Task
research_task = Task(
description=(
"Find the three closest competitors to {product}. For each, capture "
"pricing tiers, the headline promise on their homepage, and one "
"product change from the last 90 days."
),
expected_output=(
"A markdown table with one row per competitor and columns for "
"pricing tier, headline promise, recent change, and a source URL "
"in every cell. Add a note where a figure is more than 12 months old."
),
agent=researcher,
)
Notice the column list. “A table of competitor info” gives you four different tables across four runs. Named columns give you the same table every time, which matters the moment you wire this into anything automated.
Pass context so the chain actually chains
Sequential tasks forward output automatically, but I still set context explicitly. It makes the dependency visible and stops the analyst from hallucinating around a gap.
analysis_task = Task(
description=(
"Using the research table, identify three positioning gaps {product} "
"can credibly own. Flag any claim resting on a figure older than 12 months."
),
expected_output="Three gaps as bullets, each with a one-sentence rationale and a confidence level.",
agent=analyst,
context=[research_task],
)
Then assemble everything. Process.sequential runs agents in order; Process.hierarchical adds a manager agent that delegates and can re-run work it judges weak. Hierarchical sounds better on paper and costs noticeably more in tokens, because the manager reviews every step. Start sequential.
from crewai import Crew, Process
crew = Crew(
agents=[researcher, analyst, writer],
tasks=[research_task, analysis_task, writing_task],
process=Process.sequential,
verbose=True,
)
result = crew.kickoff(inputs={"product": "our inventory app for bike shops"})
print(result.raw)
Add tools only when an agent has nothing without them
Every tool is a new failure mode. A search tool can return nothing, a scraper can hit a paywall, an API can rate-limit you at task three. Add them deliberately, and give agents a fallback instruction for empty results.
When you get past search-and-scrape and need agents to touch Gmail, Slack, or a CRM, the integration layer is usually where the week disappears. Options like Composio handle authenticated tool access for agents so you’re not hand-rolling OAuth for every service, and it plugs into CrewAI’s tool interface without rewriting your crew.
Four things that break the first run
- Overlapping roles. If two agents could swap backstories without anyone noticing, they will duplicate work. Split by verb — research, judge, write — not by topic.
- Prose outputs. Anything downstream that needs structure should output markdown tables, JSON, or numbered lists. Free text gets reinterpreted, and reinterpretation drifts.
- Silent empty tool results. The agent doesn’t stop when search returns nothing; it improvises. Put “if a fact can’t be verified, write UNVERIFIED” directly in the task description.
- Runaway loops. Cap
max_iter, setmax_rpmon the Crew, and watch the verbose output for one run before you trust it on a schedule.
The verbose log is your best debugging tool. Read it once end to end. You’ll usually spot the weak link within a single run — it’s almost always a task description that was too loose, not a model that was too weak.
Running it on a schedule without bleeding money
Three practical moves before this becomes a weekly job. Pin your task outputs to a schema so downstream systems can parse them. Cache tool calls, because competitor pricing pages change monthly, not hourly. And log token usage per task — research typically dominates the bill, and trimming max_iter or narrowing the search scope usually saves more than switching models.
If you find yourself fighting type errors in the data handed between agents, it’s worth knowing the tradeoffs: PydanticAI enforces typed agent outputs in a FastAPI style, which is a different set of benefits from CrewAI’s role-based orchestration. Some teams split the work — CrewAI for the research chain, typed validation at the handoff.
And if handing Python files to non-engineers is the blocker, the visual route exists: CrewAI Studio lets you assemble and run crews through a graphical builder while producing the same underlying configuration.
One last habit worth forming. Keep the crew file in version control and re-run it after every prompt change, with the same three inputs, and diff the outputs. Agent systems drift quietly — a tweak that improves one task can gut another, and the only way to catch that is a fixed input set you run every time. Do that and your second crew takes an afternoon instead of a week.

