You have a research question. You could paste it into ChatGPT and hope for the best. But single-agent responses often miss nuance, hallucinate sources, or drift off-topic. AutoGen, Microsoft’s open-source framework for multi-agent conversations, lets you build a team of specialized AI agents that critique and improve each other’s work. In this guide, you’ll build a working research assistant with three agents: a Researcher, a Writer, and a Critic. We’ll go from installation to a finished report in six steps.
What You’ll Build (and Why It Beats a Single Prompt)
Instead of one model doing everything, you’ll create three agents with distinct roles. The Researcher gathers facts and sources. The Writer turns those facts into a clear summary. The Critic reviews the summary for accuracy, tone, and completeness. AutoGen handles the conversation loop, tool calls, and termination. The result is a report that has been checked and improved by multiple perspectives.
If you want a broader overview of what AutoGen can do, see our guide to AutoGen in practice. Here, we’re getting our hands dirty.
Step 1: Set Up Your Environment
You’ll need Python 3.10 or later. Install AutoGen with pip:
pip install pyautogen
Set your OpenAI API key as an environment variable. If you prefer local models, AutoGen supports Ollama and others, but we’ll use OpenAI’s GPT-4 for the main agents and GPT-3.5 for simpler tasks to keep costs down.
export OPENAI_API_KEY="your-key-here"
Step 2: Define Your Agents
AutoGen’s AssistantAgent is a conversational agent powered by an LLM. You define its behavior through a system message. Here’s how to create the three agents (we’ll use a shared llm_config for brevity):
import os
from autogen import AssistantAgent
llm_config = {"config_list": [{"model": "gpt-4", "api_key": os.environ["OPENAI_API_KEY"]}]}
researcher = AssistantAgent(
name="Researcher",
system_message="You are a meticulous research assistant. When given a topic, you search for recent, reliable information and return a list of key facts with sources. Be concise.",
llm_config=llm_config
)
writer = AssistantAgent(
name="Writer",
system_message="You are a clear technical writer. Take the researcher's facts and write a 200-word summary for a tech executive. Use plain language.",
llm_config=llm_config
)
critic = AssistantAgent(
name="Critic",
system_message="You are a critical editor. Review the writer's summary for accuracy, clarity, and missing details. If it's good, reply 'APPROVED'. Otherwise, explain what needs fixing.",
llm_config=llm_config
)
Crafting System Messages That Work
The system message is your main control lever. Vague instructions lead to vague output. Use these guidelines:
- Give a specific role. “You are a research assistant” is better than “You are helpful.”
- Define the output format. Specify length, structure, or tone.
- Set constraints. Limit word count, require sources, or forbid speculation.
- Include a termination trigger. The Critic’s “APPROVED” is a simple example.
Step 3: Give Your Researcher Real Tools
A researcher without search tools will make things up. AutoGen supports function calling, so you can register a Python function that performs a web search. First, create a UserProxyAgent to execute the function:
from autogen import UserProxyAgent
user_proxy = UserProxyAgent(
name="User",
human_input_mode="NEVER",
max_consecutive_auto_reply=0,
code_execution_config=False
)
Now define and register a search function:
from duckduckgo_search import DDGS
def search_web(query: str) -> str:
results = DDGS().text(query, max_results=3)
return "\n".join([f"{r['title']}: {r['body']}" for r in results])
researcher.register_for_llm(name="search_web", description="Search the web for current information")(search_web)
user_proxy.register_for_execution(name="search_web")(search_web)
If writing tool wrappers feels tedious, tools like Composio can give your agents real-world integrations without the integration headache. But for this tutorial, a simple search function is enough.
Step 4: Orchestrate the Conversation
Now you need a manager to coordinate the agents. AutoGen’s GroupChat and GroupChatManager handle turn-taking.
from autogen import GroupChat, GroupChatManager
groupchat = GroupChat(
agents=[user_proxy, researcher, writer, critic],
messages=[],
max_round=12,
speaker_selection_method="auto"
)
manager = GroupChatManager(groupchat=groupchat, llm_config=llm_config)
The speaker_selection_method can be “auto” (the manager picks the next speaker) or “round_robin”. Auto is more flexible but can be less predictable. Round-robin is good for a fixed workflow. Start with auto and see how it goes.
Step 5: Add a Termination Condition
Without a stop condition, agents can chat forever. Use is_termination_msg to end when the Critic approves:
def is_termination_msg(msg):
return "APPROVED" in msg.get("content", "")
groupchat = GroupChat(
agents=[user_proxy, researcher, writer, critic],
messages=[],
max_round=12,
speaker_selection_method="auto",
is_termination_msg=is_termination_msg
)
You can also set max_round as a hard stop. A good default is 10-15 rounds for a research task.
Step 6: Run It and Iterate
Give the team a task:
task = "Research the latest advances in solid-state batteries and write a 200-word summary for a tech executive. Include at least two sources."
user_proxy.initiate_chat(manager, message=task)
The conversation will flow something like this: the Researcher searches and returns facts, the Writer drafts a summary, the Critic reviews it, and if changes are needed, the Writer revises. The loop continues until “APPROVED” appears or you hit the round limit. On a typical run with GPT-4, expect 30-60 seconds and a few cents in API costs.
Keep a log of each run. You’ll quickly spot patterns: maybe the Researcher needs a stricter prompt, or the Critic is too lenient. Tweak the system messages and run again.
Leveling Up: Caching, Cost Control, and Alternatives
AutoGen has a built-in caching mechanism to avoid repeated LLM calls for identical prompts. Enable it with cache_seed=42 in your llm_config. For cost control, mix models: use GPT-4 for the Writer and Critic, but GPT-3.5 for the Researcher’s initial search summarization. You can also limit the number of search calls per turn.
If you prefer a lighter, code-first approach, SmolAgents strips away much of the conversation overhead. For enterprise deployment with built-in monitoring, IBM BeeAI is worth exploring. And if you’re not a coder, our guide to making AI agents in 2025 covers no-code and low-code options.
Common Pitfalls and How to Avoid Them
- Agents looping. They repeat the same message. Fix: add a termination condition, reduce
max_round, or make the Critic’s instructions more decisive. - Ignoring instructions. The LLM forgets its role. Fix: put the most important constraints in the system message and repeat them in the task prompt.
- API errors. Rate limits or timeouts. Fix: add retry logic with exponential backoff in your
llm_config. - High costs. Long conversations with GPT-4 add up. Fix: use cheaper models for simple turns, enable caching, and set a hard round limit.
Extending the Pattern: Add a Code Executor
Once your research assistant works, you can add a fourth agent that executes code. Give a UserProxyAgent the code_execution_config and let it run Python for data analysis, chart generation, or calculations. The Researcher can pass structured data to the Writer, and the Code Executor can verify numbers. This pattern scales to financial analysis, scientific literature reviews, and competitive intelligence. The core loop stays the same: specialized agents, clear roles, and a termination condition that tells everyone when the job is done.

