SyncAI.news, a Varaisys broadcasting
Building an AI-powered contract intelligence platform with Amazon Quick and Amazon Bedrock AgentCore
KM

Konala McGrath

· 12 min read

EngineeringAWS Machine Learning Blog

Building an AI-powered contract intelligence platform with Amazon Quick and Amazon Bedrock AgentCore

You’re a director of contracting, responsible for hundreds, maybe thousands, of vendor contracts. Each one is packed with critical data: contract values, expiration dates, signing status, and key contacts. With that information locked inside PDFs, you and your team spend hours manually extracting it, maintaining spreadsheets, and fielding the same recurring questions: “Which vendor are we spending the most with?” or “Which contracts are about to expire?”

So you turn to AI chat tools and enterprise Q&A solutions. You upload a contract, ask questions in natural language, and for a single document, or even a handful, it genuinely works well. The real challenge shows up at scale, across hundreds of contracts.

Most of these tools rely on a technique called Retrieval Augmented Generation (RAG). RAG breaks long text into chunks and indexes each one individually. During a search, RAG retrieves only the top-k chunks of text most relevant to a question, an approach more widely known as semantic search. That’s fine when the answer lives in one place. But when a question spans every contract, like total exposure or upcoming renewals, semantic search falls short. The answer requires aggregation across the full dataset, not only a handful of chunks. A better prompt won’t fix this. A different architecture will.

In this post, we share the architecture behind a contract intelligence platform on AWS. It extracts contract data with AI agents, verifies accuracy with a multi-model approach, and delivers instant answers through embedded analytics and natural language querying. Everything runs from a single web application.

The challenge: When AI chat tools hit their limit

Say your portfolio holds 250 vendor contracts, and leadership keeps asking the same four questions:

  • Total contract value across the portfolio.
  • Which contracts have already expired.
  • The most expensive agreement in the portfolio.
  • The most recently signed deal.

These are basic questions, but the answers are buried. Each of your 250 contracts runs 10–20 pages, so that’s up to 5,000 pages of unstructured data. To answer those questions, an analyst opens each PDF, searches for the fields that matter, and copies the values into a spreadsheet, then repeats the process 250 times. That’s more than a week of effort to clear the backlog, before a single new contract arrives. Every follow-up question from leadership means another round of filtering and analysis. Each round risks working from stale data or breaking hardcoded formulas on the new data.

So you upload the portfolio to one of the leading AI chat tools and ask for the total value. The answer comes back fast, confident, and wrong.

Where RAG fits, and where aggregation is needed

That wrong answer wasn’t a fluke. It’s baked into how these tools are built. Internally they chunk each document into vectors and store them in a knowledge base. When you ask a question, they pull back the top-k chunks that best match it, then build an answer using only those chunks as context.

For a targeted lookup, that’s exactly what you want. Ask “What are the payment terms in the AnyCompany contract?” and the right chunk surfaces with a clean answer. A portfolio question works against that mechanism instead of with it. When you ask, “What’s the total contract value across all 250 contracts?”, the system still returns only the top-k chunks and totals only those. It can’t sum, count, or compare across the portfolio because the portfolio never lands in front of the model. It isn’t a flaw in any one tool but in how RAG itself works.

The key insight

The answer is structured extraction, not better retrieval. To answer aggregation questions across hundreds of documents, you extract the key fields into a database, then query them with analytics tools built for exactly that. AI handles the unstructured-to-structured conversion at scale, and the database is used for math. You keep the original documents in a knowledge base for the single-contract questions it handles well. The result is a system that answers both kinds of questions: portfolio-wide aggregates and precise single-document lookups.

Solution overview

The platform is a React web application on AWS that turns a portfolio of contract PDFs into structured, queryable data. AI agents extract the key fields from each contract and independently verify them, with Amazon Textract settling any signature disagreements. Users then explore the results through embedded dashboards and a natural language chat agent, getting both portfolio-wide answers and single-contract details.

The pipeline works as follows:

  1. You store contract PDFs in an Amazon Simple Storage Service (Amazon S3) bucket, which triggers an automated processing pipeline.
  2. An AI extraction agent (powered by Claude Sonnet series of models) reads the PDF and extracts eight key fields, each with a confidence score.
  3. A separate AI verification agent (powered by lighter Claude Haiku model) independently reads the same PDF and verifies the extraction.
  4. If the two agents disagree on signature detection, Amazon Textract provides a deterministic tiebreaker using computer vision.
  5. The system stores verified results in Amazon Aurora PostgreSQL.
  6. Users see real-time pipeline status over WebSocket and can immediately query data through embedded dashboards and a natural language chat agent.

The entire pipeline can process a contract in seconds under typical conditions, and the serverless architecture is designed to help scale to handle many contracts in parallel. Here is the reference architecture for this solution.

Figure 1: Contract intelligence pipeline, from ingestion to query

Note: This post references the language models available when the solution was built. Model options change quickly, use the models available to you and re-test the solution before relying on the results.

Inside the AI extraction pipeline

The pipeline turns each uploaded contract into verified, structured data. It starts with two agents, one that extracts the fields and one that independently checks them, both running on fully managed infrastructure. The following section covers how they work together and the models behind them.

Dual-model verification with Strands Agents on Amazon Bedrock AgentCore

The extraction and verification agents are built with the Strands Agent SDK and deployed on AgentCore runtime, a capability of Amazon Bedrock AgentCore. The Strands Agent SDK is an open-source, model-driven framework for building autonomous agents. AgentCore handles serverless hosting, automatic scaling, and session isolation, so there’s no infrastructure to manage for the AI components.

With Policy in Amazon Bedrock AgentCore, you can define and enforce security controls as a protective boundary around agent operations. This securely handles sensitive data such as pricing terms, financial commitments, and vendor relationships. Cedar-based policies exist outside agent code, and you can validate them with automated reasoning before enforcement. Agents and users only access the contracts they are authorized to see.

We deliberately chose two different foundation models for extraction and verification:

  • Claude Sonnet 4.6 for extraction: Advanced document understanding capabilities, reads the full PDF natively (no optical character recognition (OCR) preprocessing needed), and returns structured JSON with confidence scores for each field.
  • Claude Haiku 4.5 for verification: Fast, cost-effective, and provides an independent perspective. Because it’s a different model with different training, it catches errors that re-running the same model would miss.

For model availability by Region, refer to Supported models by AWS Region in Amazon Bedrock.

This dual-model pattern is key to high accuracy on contract data, where mistakes are costly. A single model might hallucinate a value with high confidence. Two independent models disagreeing is a strong signal that human review is needed.

Amazon Textract: The signature detection tiebreaker

During testing, we found a revealing failure mode. The verification model would sometimes flag a contract as “signed” with 95–100% confidence when no signature existed, disagreeing with the extractor, which had correctly read it as unsigned. The model misread empty signature blocks, treating the presence of a signature field (“Signature: __________”) as proof of a signature.

Rather than accepting this hallucination or writing elaborate prompts to work around it, we added Amazon Textract computer-vision-based signature detection as an architectural tiebreaker. Amazon Textract uses visual analysis, not language understanding, to detect actual handwritten or digital signatures on the page. It runs only when the two models disagree on the is_signed field, which keeps costs minimal while catching the false positives.

This illustrates an important pattern: use large language models (LLMs) for what they excel at (document comprehension, field extraction, contextual understanding) and deterministic services for what they struggle with (visual element detection, precise counting, mathematical operations).

Data-driven model selection

The pipeline’s accuracy depends most on the model that handles extraction, so we tested several extractor and verifier combinations before settling on one. Using our 20-contract sample dataset, we hand-labeled all 8 fields (160 values total) to establish ground truth, then measured how closely each combination matched it.

The evaluation surfaced two patterns. Extraction is the higher-impact step. A more capable extraction model held accuracy up even when paired with a lighter verifier, while a lighter extraction model brought accuracy down regardless of the verifier. More capability wasn’t always worth the cost. The strongest models didn’t meaningfully outperform a capable, lower-cost verifier on our data.

Amazon Bedrock gives you a broad choice of models, and the right combination depends on your data. This was a small, directional test on 20 contracts, not an exhaustive evaluation, and results will vary with contract format and field complexity. We strongly recommend running your own evaluation against your benchmarks and success criteria, using the models available to you at build time.

Querying at scale with Amazon Quick

Extracting and verifying the data solves only half the problem. People still need to ask questions of it, from broad portfolio totals to the details of a single contract. Amazon Quick connects both the structured records and the original documents behind a single interface, so anyone can ask either kind of question. Within Amazon Quick, the Amazon Quick Sight capability powers the dashboards, while natural language querying and chat handle plain-language questions.

Bridging structured and unstructured

With contract data now extracted into a PostgreSQL database, we connected it to Amazon Quick so users can explore it through visual dashboards and natural language querying. The result is a single interface that answers both types of question:

  • Aggregation queries (through structured data and Topics): “What’s our total contract value?” → “$50 million across 20 contracts.”
  • Document-specific queries (through the knowledge base): “What are the payment terms in the AnyCompany contract?” → the exact clause, pulled from the source PDF.

Embedded dashboards

The Amazon Quick Sight dashboards are embedded directly in the React application using the Amazon Quick Sight Embedding SDK. Without leaving the application, users get the contract analytics they need without requiring exports or separate business intelligence (BI) tools. Rather than caching in SPICE (Super-fast, Parallel, In-memory Calculation Engine) and incurring extra costs, we query the data directly, so the dashboards are always live. The moment a contract finishes processing, it shows up.

The dashboard includes several views:

  • Overview: Key performance indicators (KPIs) such as total contracts and total value, a signed and unsigned pie chart, and a full extraction table with filters.
  • Details: Confidence scores from both the extractor and verifier agents, with field-by-field match status.
  • Costs: A service-level cost breakdown that highlights the serverless architecture.

The chat agent: Natural language meets analytics

With the embedded chat agent in Amazon Quick, you can ask questions about your contract portfolio in natural language. Because it connects to both a Topic (for structured database queries) and a knowledge base (for document-level search), you can ask the full spectrum of questions:

  • “How many contracts do I have and what’s the total value?” → 20 contracts, $50 million.
  • “How many contracts are expired, and what’s their total value?” → 3 expired, $3.7 million.
  • “How many are signed versus unsigned?” → 17 signed, 3 unsigned.
  • “What are my top 3 contracts to prioritize?” → the unsigned contracts nearing expiration, ranked by value.

Structured analytics and document Q&A in a single chat: That’s the payoff, aggregation and document lookups answered from one interface.

Real-time pipeline transparency

One design decision that improved the user experience was adding WebSocket-based real-time status updates. When contracts enter the pipeline, the React frontend displays each processing stage as it happens. This turns an opaque upload-and-wait experience into a transparent process where users see exactly which service is working on their contracts. It builds trust and makes the multi-model flow clear.

Cost analysis

The solution is primarily serverless, with the bulk of ongoing cost in Amazon Quick licensing rather than AI inference:

Service Monthly Cost (at 1,000 contracts/month)
Amazon Quick (1 Enterprise user + infrastructure fee)* $290
Amazon Aurora Serverless v2 $44
Amazon Bedrock (Sonnet extraction + Haiku verification) $12
Amazon Textract (signature tiebreaker) $2
AWS Lambda, Amazon API Gateway, Amazon S3, Amazon CloudFront < $1
Total ~$349/month

* Pricing is based on the date the blog was published. Visit the Amazon Quick Pricing page for latest accurate information.

At $0.014 per contract for AI inference (Amazon Bedrock plus Amazon Textract), the variable cost scales linearly and stays negligible next to the fixed analytics licensing. The economic savings compared to manual work are clear.

Key takeaways

  1. Extract, then query: For aggregation use cases across many documents, extract key fields into a structured database. RAG excels at finding information in documents. Databases excel at aggregating across them. Use both.
  2. Verify with independence: Using two different foundation models for extraction and verification catches errors that a single model would miss. The cost of a second, cheaper model is minimal compared to the accuracy gain on high-stakes data.
  3. Use deterministic services for deterministic tasks: When LLMs hallucinate (signature detection, precise counting), bring in purpose-built services like Amazon Textract rather than adding a third LLM. Match the tool to the task.
  4. Make the AI visible: Real-time pipeline status builds user trust and makes multi-agent architectures tangible to stakeholders. Transparency turns an opaque AI pipeline into a comprehensible workflow.
  5. This pattern scales beyond contracts: Domains with large volumes of unstructured documents requiring portfolio-level queries benefit from this approach: procurement, legal compliance, insurance claims, HR onboarding documents, real estate leases.

Conclusion

We started with a director who couldn’t get a straight answer about his own contracts. By designing for structured extraction first and pairing analytics tools with AI agents, this architecture turns an unmanageable contract portfolio into one that answers questions in seconds. That holds whether the question is a total across hundreds of contracts or the payment terms in one. The pattern is straightforward: use LLMs for text-based analysis and a database for numerical queries. Where two models disagree in ways a deterministic service can’t settle, the next step is to route the contract to a human reviewer instead of auto-resolving. This keeps people in the loop exactly where judgment matters most.

To get started with the services used in this architecture, explore the following resources:

  • Amazon Bedrock AgentCore documentation — for deploying and managing AI agents at scale.
  • Strands Agents SDK on GitHub — for building multi-model agent workflows.
  • Amazon Quick Sight Embedding SDK — for integrating dashboards and natural language querying into your applications.
  • Amazon Textract HumanLoopConfig documentation — for document analysis and signature detection.
  • Amazon Quick Human-in-the-loop task center — for adding human judgment at critical points in your automated processes.

About the authors

Konala McGrath

Konala is a Solutions Architect at AWS. He supports digital native businesses across multiple industries and helps them build highly scalable, cost-optimized cloud solutions. In his free time, he enjoys spending time with his family and friends, and watching or playing sports.

Alberto Alonso

Alberto is a Specialist Solutions Architect at Amazon Web Services. He focuses on generative AI and how it can be applied to business challenges.

Hugo Tse

Hugo is a Solutions Architect at AWS supporting digital native businesses. He strives to help customers use cloud technologies to tackle problems and create new business opportunities, especially in the domains of generative AI and storage.

Nitish Chaudhari

Nitish is a Senior Solutions Architect at AWS. He specializes in using Amazon’s Generative AI offerings for building productivity solutions for engineers, knowledge workers and data analytics teams and scaling adoption of those solutions across organizations.

Original source

This story was published by AWS Machine Learning Blog and written by Konala McGrath. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on aws.amazon.com

Similar News