
Manideep Reddy Gillela
· 14 min read
Agentic retrieval with LangChain and Amazon Bedrock Knowledge Bases
When a user asks the support assistant, a Retrieval Augmented Generation (RAG) application built with LangChain to compare two products across three dimensions, they’re effectively posing six questions simultaneously. Similarity search uses a single query vector to encapsulate all the intents. The retriever then generates the best approximation of the average of those intents. The resulting answer comes back concise. The search executes without errors. The relevance scores look reasonable. Yet the retrieved chunks, while topically relevant, only cover a fraction of what the question actually asked.
In this post, we showcase a RAG application on Amazon Bedrock Managed Knowledge Base with LangChain. We run the same multi-part question through standard and agentic retrieval, and read the trace events to see the plan the model produced. We also cover what the two retrieval paths cost and when the cheaper one is the right choice.
Agentic retrieval is available on Amazon Bedrock Managed Knowledge Base. Instead of one search, Amazon Bedrock Managed Knowledge Base plans the retrieval. It breaks the question into sub-queries, runs them, judges whether it has enough evidence, and searches again if it doesn’t. The langchain-aws package exposes both agentic and standard retrieval, so you can use either from a LangChain application.
Solution overview
Amazon Bedrock Managed Knowledge Base, the fully managed RAG capability in Amazon Bedrock, removes the self-managed vector store, embeddings, and re-ranking models from the RAG architecture. You configure a data source, and Amazon Bedrock Managed Knowledge Bases handles chunking, embedding, storage, and retrieval. This walkthrough uses Amazon Simple Storage Service (Amazon S3).
Amazon Bedrock Managed Knowledge Bases provides two APIs. We briefly discuss those differences in this post. The Retrieve API runs one hybrid search and returns scored chunks. The AgenticRetrieveStream API runs a planning loop and streams the steps back to you as trace events. In the langchain-aws package, the first is a standard LangChain retriever you can drop into a chain. The second is a function retrieval directly from a knowledge base.
The following diagram shows the solution architecture. The application queries Amazon Bedrock Knowledge Bases using either the Retrieve API (standard, single-shot) or the AgenticRetrieveStream API (multi-step planning loop). Both paths return document chunks from the knowledge base, which the application then uses to generate a grounded response.
Figure 1: Solution architecture for querying Amazon Bedrock Knowledge Bases with the Retrieve and AgenticRetrieveStream APIs
Implementation walkthrough
The following sections walk you through creating a knowledge base, querying it with both retrieval methods, and reading the trace events the agentic planner produces.
Prerequisites
To follow along you need:
- An AWS account with access to Amazon Bedrock in a Region where Amazon Bedrock Managed Knowledge Bases and agentic retrieval are available. This walkthrough uses the US East (N. Virginia) Region (
us-east-1), and the code assumes it throughout. Check the AWS documentation for other Regional availability and support. - Two AWS Identity and Access Management (IAM) identities, described in the next section: a service role the knowledge base assumes, and permissions on the identity you call the APIs from.
- Python 3.12 or later.
- An S3 bucket holding the sample documents. The corpus needs several documents that cover overlapping topics so that a comparative question has somewhere to go. A single flat document cannot demonstrate query planning.
Install the packages. The Boto3 version matters: agentic_retrieve_stream did not exist before 1.43.32.
langchain-aws>=1.6.3
langchain>=1.0
boto3>=1.43.32
Permissions
Two identities are involved and separating them is worth doing deliberately. The knowledge base assumes a service role to read your documents and call the embedding model. Your application uses an AWS Security Token Service (AWS STS) caller identity to query. Neither needs the other’s permissions.
Amazon Bedrock creates the service role for you if you let it. To supply your own, give it a trust policy that lets Amazon Bedrock assume it. Scope it with aws:SourceAccount and aws:SourceArn so that another account can’t use it as a confused deputy:
{
"Version": "2012-10-17",
"Statement": [{
"Effect": "Allow",
"Principal": {"Service": "bedrock.amazonaws.com"},
"Action": "sts:AssumeRole",
"Condition": {
"StringEquals": {"aws:SourceAccount": "111122223333"},
"ArnLike": {
"aws:SourceArn": "arn:aws:bedrock:us-east-1:111122223333:knowledge-base/*"
}
}
}]
}
The service role also needs s3:ListBucket on your bucket and s3:GetObject on its contents, both conditioned on aws:ResourceAccount. Scope the knowledge-base/* wildcard character down to specific knowledge base IDs after you have created them.
The AWS STS caller identity needs a different set. bedrock:AgenticRetrieveStream and bedrock:InvokeModelWithResponseStream can’t be scoped to a knowledge base Amazon Resource Name (ARN). bedrock:Retrieve and bedrock:GetDocumentContent can:
{
"Version": "2012-10-17",
"Statement": [
{
"Sid": "AgenticRetrievalAndPlannerModel",
"Effect": "Allow",
"Action": [
"bedrock:AgenticRetrieveStream",
"bedrock:InvokeModelWithResponseStream"
],
"Resource": "*"
},
{
"Sid": "RetrieveAndFullDocumentExpansion",
"Effect": "Allow",
"Action": ["bedrock:Retrieve", "bedrock:GetDocumentContent"],
"Resource": "arn:aws:bedrock:<region>:111122223333:knowledge-base/<knowledge-base-id>"
},
{
"Sid": "GenerateAnswersInTheChains",
"Effect": "Allow",
"Action": ["bedrock:InvokeModel", "bedrock:Converse", "bedrock:ConverseStream"],
"Resource": "*"
}
]
}
bedrock:GetDocumentContent is often overlooked. Agentic retrieval calls it when a FullDocumentExpansion step decides a passage lacks the context to answer. A policy with only bedrock:Retrieve works until the planner reaches for a whole document and then fails partway through a query.
To create and manage the knowledge base itself, the calling role additionally needs bedrock:CreateKnowledgeBase on *, and the GetKnowledgeBase, UpdateKnowledgeBase, DeleteKnowledgeBase, StartIngestionJob, GetIngestionJob, and ListIngestionJobs actions on knowledge-base/*. If you’re using guardrails, add bedrock:GetGuardrail and bedrock:ApplyGuardrail.
Running this walkthrough might incur costs for document storage and ingestion in the knowledge base, retrieval calls, and foundation model (FM) inference.
Delete the resources when you complete this experiment.
Creating and populating the knowledge base
Create the knowledge base with a managedKnowledgeBaseConfiguration. Setting embeddingModelType to MANAGED uses the service-managed embedding model.
import boto3
import os
REGION = os.environ["AWS_REGION"]
bedrock_agent = boto3.client("bedrock-agent", region_name=REGION)
response = bedrock_agent.create_knowledge_base(
name=KB_NAME,
roleArn=KB_ROLE_ARN,
knowledgeBaseConfiguration={
"type": "MANAGED",
"managedKnowledgeBaseConfiguration": {
"embeddingModelType": "MANAGED",
},
},
)
KB_ID = response["knowledgeBase"]["knowledgeBaseId"]
There’s no storageConfiguration in that request. For a self-managed knowledge base you would pass one describing your vector store. Amazon Bedrock Managed Knowledge Base does not take one, which is the clearest signal in the API that Amazon Bedrock owns the storage layer.
Attach the S3 bucket as a data source, then start an ingestion job. Ingestion is asynchronous, so poll until the job reaches a terminal state rather than sleeping for a fixed interval and hoping.
import time
SUCCESS_STATES = frozenset({"COMPLETE"})
FAILURE_STATES = frozenset({"FAILED", "STOPPED"})
def wait_for_ingestion(kb_id, ds_id, job_id, timeout_s=1800):
"""Poll an ingestion job until it reaches a terminal state."""
deadline = time.time() + timeout_s
while time.time() < deadline:
job = bedrock_agent.get_ingestion_job(
knowledgeBaseId=kb_id,
dataSourceId=ds_id,
ingestionJobId=job_id,
)["ingestionJob"]
status = job["status"]
if status in SUCCESS_STATES:
return job
if status in FAILURE_STATES:
reasons = job.get("failureReasons") or ["no reason reported"]
raise RuntimeError(f"Ingestion job {job_id} finished as {status}: " + "; ".join(reasons))
time.sleep(15)
raise TimeoutError(f"Ingestion job {job_id} did not finish in {timeout_s}s")
The full data source configuration and error handling are in the sample repository.
Querying with the LangChain retriever
AmazonKnowledgeBasesRetriever wraps the Retrieve API and behaves like any other LangChain retriever. For Amazon Bedrock Managed Knowledge Bases, pass managedSearchConfiguration. This is the part that trips people up: vectorSearchConfiguration is the previous path for knowledge bases where you run your own vector store. It is what most existing examples show.
from langchain_aws.retrievers import AmazonKnowledgeBasesRetriever
SIMPLE_QUERY = "What is the restore time objective for the checkout service?"
retriever = AmazonKnowledgeBasesRetriever(
knowledge_base_id=KB_ID,
region_name=REGION,
retrieval_config={
"managedSearchConfiguration": {
"numberOfResults": 5,
}
},
)
docs = retriever.invoke(SIMPLE_QUERY)
Each result comes back as a LangChain Document. The relevance score is in metadata["score"], and the source document’s own metadata is under metadata["source_metadata"], renamed so it does not collide. If you want to drop low-confidence results, set min_score_confidence on the retriever instead of filtering afterward.
For a question with one clear intent, this is the right tool. It is one call. The latency is the lowest of the two options, and you keep full control of how the answer gets generated. Most of the queries a production assistant sees are this shape, and reaching for a planning loop to answer them wastes money and time.
Where single-shot retrieval runs out
Now give the same retriever a question with several parts:
COMPLEX_QUERY = (
"Compare the checkout and inventory services across on-call escalation, backup and "
"restore targets, and deployment rollback procedure. Where do they differ?"
)
docs = retriever.invoke(COMPLEX_QUERY)
Five chunks come back, ranked by hybrid score against one embedding of that whole question.
That question contains six intents: two services across three dimensions. Scoring the retrieved text for evidence of each one gives a concrete measure of what a single embedding recovers.
| numberOfResults | Chunks | Share of corpus | Sub-intents covered | Missing |
| 5 | 5 | 10% | 4 of 6 | checkout on-call, inventory restore |
| 10 | 10 | 19% | 6 of 6 | none |
At five results, one embedding standing for six intents misses two of them. At ten it covers all six, with visible waste: two sub-intents are covered twice and one chunk carries none.
The retriever did its job. The limitation is structural: one vector cannot represent six intents, and there is no step in the process that asks whether the returned evidence is enough to answer the question.
Running agentic retrieval
Agentic retrieval is not a LangChain retriever, but a feature of Amazon Bedrock Managed Knowledge Bases. The langchain-aws package exposes it as a standalone function, agentic_retrieve, because the underlying API streams its results and doesn’t fit the synchronous BaseRetriever interface. There’s no flag on AmazonKnowledgeBasesRetriever that switches it on.
from langchain_aws.retrievers.bedrock import agentic_retrieve
result = agentic_retrieve(
knowledge_base_id=KB_ID,
query=COMPLEX_QUERY,
region_name=REGION,
generate_response=True,
number_of_results=10,
)
print(result["generatedResponse"]["answer"])
With generate_response=True, the service returns a grounded answer and citations alongside the retrieved chunks, so you get an answer without wiring up a separate model call. The function works only against Amazon Bedrock Managed Knowledge Base.
Internally, the service plans, retrieves, evaluates whether the evidence is sufficient, and iterates if it isn’t. The helper hides all of that and hands back the final chunks, which is convenient and means you can’t see the plan.
Reading the trace events
To watch the model decompose the question, call agentic_retrieve_stream on the bedrock-agent-runtime client directly. This is the one place in this walkthrough where we step around langchain-aws, because the helper discards trace events and doesn’t expose maxAgentIteration or a custom planner model.
runtime = boto3.client("bedrock-agent-runtime", region_name=REGION)
response = runtime.agentic_retrieve_stream(
messages=[{"role": "user", "content": {"text": COMPLEX_QUERY}}],
retrievers=[{
"configuration": {
"knowledgeBase": {
"knowledgeBaseId": KB_ID,
"retrievalOverrides": {"maxNumberOfResults": 10},
}
}
}],
agenticRetrieveConfiguration={
"foundationModelType": "MANAGED",
"rerankingModelType": "MANAGED",
# 5 is the API default. Below 4 the planner stops decomposing entirely.
"maxAgentIteration": 5,
},
generateResponse=False,
)
for event in response["stream"]:
if "traceEvent" in event:
attrs = event["traceEvent"]["attributes"]
print(f"{attrs.get('step')}: {attrs.get('status')}")
for action in attrs.get("actions", []) or []:
if "retrieve" in action:
query = action["retrieve"].get("inputQuery", {}).get("text", "")
print(f" sub-query: {query}")
elif "result" in event:
for chunk in event["result"].get("results", []):
print(chunk.get("content", {}).get("text", "")[:120])
Set generateResponse to False when you only want the retrieval behavior. The API generates a grounded answer by default, which costs an extra model call you may not need while you are inspecting the plan.
A trace event’s step tells you where the planner is. SpeculativeRetrieval runs before the first plan to cut latency and doesn’t count against your iteration budget. Planning is where the model reads the question and prior results and emits sub-queries. Retrieval fires once per sub-query. FullDocumentExpansion appears when the model decides a passage lacks the context to answer and pulls the whole document instead. Each carries a status of IN_PROGRESS, SUCCEEDED, or FAILED, plus a human-readable message.
The final chunks arrive separately. The result is its own event type rather than a fifth step, and it holds the deduplicated chunks from every iteration along with the grounded answer when response generation is on. Branch on the event key, as the preceding loop demonstrates, rather than expect a terminal step value.
The sub-query text is the part worth logging. It sits in attributes.actions[].retrieve.inputQuery.text, not in the top-level trace fields, so a handler that reads only step and status shows you that planning happened without showing you what it decided.
The following diagram shows the agentic retrieval planning loop, including the speculative retrieval, planning, sub-query retrieval, evaluation, and optional re-planning steps.
Figure 2: Steps in the agentic retrieval planning loop
Two details are worth knowing before you build on this. Deduplication applies only to the result event, so a chunk retrieved by three sub-queries appears once at the end but three times across the traces.
The second is about scores. A Retrieve response gives each chunk a typed score field holding its relevance to the query. Agentic retrieval results carry content, metadata, and sourceRetriever, with no equivalent typed field. Code that reads result["score"] after switching APIs gets nothing. If you rank or filter relevance, plan for that difference.
In production, use Amazon Bedrock Guardrails to enforce content policies and grounding checks on generated responses. Both retrieval paths support guardrails. Agentic retrieval supports guardrails through policyConfiguration.bedrockGuardrailConfiguration rather than the guardrail_config argument the LangChain retriever takes, and supports BLOCK mode only. If you rely on MASK mode, that is a reason to stay on the Retrieve API.
maxAgentIteration accepts two through ten and defaults to five. Leave it at the default. At two or three the planner runs one cycle, emits no sub-queries, and returns what the speculative retrieval step already found. This is single-shot behavior at the agentic price. Decomposition begins at four. The planner often stops early when it judges the evidence sufficient, so the ceiling is a bound rather than a target.
Comparing the two retrieval paths
For context on how this behaves at scale, AWS evaluated agentic retrieval on MuSiQue, a public multi-hop benchmark. The evaluation showed improved recall over single-shot retrieval, with the largest gains on the hardest questions. Single-hop questions saw gains under five points. That last figure matches the shape of the trade-off: decomposition helps when there is something to decompose.
Building the RAG chain
For the standard retriever, the usual LangChain Expression Language (LCEL) composition works directly:
from langchain_aws import ChatBedrockConverse
from langchain_core.output_parsers import StrOutputParser
from langchain_core.prompts import ChatPromptTemplate
from langchain_core.runnables import RunnablePassthrough
def format_docs(docs):
return "\n\n".join(doc.page_content for doc in docs)
llm = ChatBedrockConverse(model=MODEL_ID, region_name=REGION)
prompt = ChatPromptTemplate.from_template(PROMPT_TEMPLATE)
chain = (
{"context": retriever | format_docs, "question": RunnablePassthrough()}
| prompt
| llm
| StrOutputParser()
)
format_docs matters more than it looks. Passing Document objects straight into a prompt renders their repr, and the model gets metadata noise mixed into the context.
To put agentic retrieval in the same position, wrap it in a RunnableLambda, since it is a function rather than a retriever:
from langchain_core.runnables import RunnableLambda
def agentic_context(question: str) -> str:
result = agentic_retrieve(
knowledge_base_id=KB_ID,
query=question,
region_name=REGION,
number_of_results=10,
)
return "\n\n".join(
item.get("content", {}).get("text", "")
for item in result.get("results", [])
)
agentic_chain = (
{"context": RunnableLambda(agentic_context), "question": RunnablePassthrough()}
| prompt
| llm
| StrOutputParser()
)
Note that generate_response is off here. The service can generate the answer itself, but inside a chain you usually want your own prompt and model, so you take the chunks and generate downstream. Use the service generation when you want one call and less code, and the wrapped version when the prompt is yours to control.
Choosing between standard and agentic retrieval
Use Retrieve for short, well-scoped questions. It is cheaper, faster, works against self-managed knowledge bases, and returns scores in the results. Most production traffic looks like this.
Use AgenticRetrieveStream when questions are multi-part, comparative, or exploratory, or when the evidence spans more than one knowledge base. It registers up to five knowledge bases in one request and routes sub-queries using a natural-language description you attach to each. The other API cannot do this at all. It costs more per call, makes several model invocations, and has the higher latency of the two.
Routing on query shape rather than picking one for everything is the pattern we recommend. A classifier or a heuristic on the question can send most traffic down the cheap path and reserve the planner for questions that need it.
Clean up resources
Delete the knowledge base, its data source, the S3 objects and bucket, and the IAM role that you created. A knowledge base with documents in it continues to incur storage charges.
bedrock_agent.delete_data_source(knowledgeBaseId=KB_ID, dataSourceId=DS_ID)
bedrock_agent.delete_knowledge_base(knowledgeBaseId=KB_ID)
The repository includes a cleanup script that also empties the bucket and removes the role.
Conclusion
We showed how to build a RAG application on Amazon Bedrock Knowledge Bases with LangChain, and how agentic retrieval handles multi-part questions that single-shot retrieval answers poorly. We also showed the friction in the current integration. Agentic retrieval is a function rather than a LangChain retriever, so it needs a RunnableLambda to sit in a chain. The trace events that show the query plan require a direct boto3 call.
Agentic retrieval trades higher per-call cost for improved recall on multi-hop questions, using a built-in model for query planning. The next useful step is measuring your own query mix before you route everything through a planner.
To get started, see the Amazon Bedrock Knowledge Bases documentation and the accompanying sample code. For help applying this to your own workload, contact your AWS account team.
About the authors
Manideep Reddy Gillela
Manideep is a Delivery Consultant, Cloud Infrastructure Architect at Amazon Web Services, where he helps enterprise customers design scalable, secure, and cost-effective cloud solutions. With more than six years of experience spanning cloud architecture, AI/ML, and generative AI on AWS, he partners with organizations to accelerate digital transformation. Outside of work, Manideep enjoys traveling, swimming, and playing recreational sports.
Luis Felipe Florez Leano
Luis is a Solutions Architect on the Americas GenAI Partner Solutions Architecture team at AWS. He works with AWS Partners across the Americas to help them design, build, and scale generative AI solutions on AWS. His focus is on practical implementations using Amazon Bedrock and helping organizations navigate the technical and business opportunities of generative AI.
Sushma Sunkollu Nagaraj
Sushma is a Partner Solutions Architect at Amazon Web Services with over five years of experience helping partners and customers build secure, scalable cloud solutions. Specializing in DevOps and infrastructure automation, she collaborates with strategic partners to design AWS-optimized architectures, lead technical workshops, and deliver high-impact proofs-of-concept. Her expertise extends into AI/ML, where she supports customers in building intelligent applications using AWS AI services. She is passionate about simplifying complexity and enabling innovation at scale.
Satya Pattanaik
Satya is a Senior Solutions Architect at AWS. In this role, he assists Independent Software Vendors (ISVs) in developing scalable and resilient applications on the AWS Cloud. Prior to joining AWS, he made a substantial impact in driving growth and cloud adaptation for large organizations. Outside of his professional endeavors, Satya is passionate about BBQ. He continuously learns new techniques to create flavorful dishes and experiments with various recipes.
Original source
This story was published by AWS Machine Learning Blog and written by Manideep Reddy Gillela. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on aws.amazon.com


