EXECUTIVE OVERVIEW
In the rapidly evolving landscape of generative artificial intelligence and autonomous workflows, developers face a deceptive paradox: a system can achieve near-perfect evaluation metrics while harboring critical, silent failures in production. This phenomenon—often masked by automated LLM-as-a-judge scoring frameworks—poses significant risks as enterprises increasingly transition from experimental text-generation chatbots to autonomous agents endowed with tools, databases, and real-world system access.
This reality became starkly evident during a recent architectural build for the Udacity and AWS Agent Engineer Nanodegree program. Tasked with developing an advanced, context-aware customer-support chatbot for a fictional e-commerce platform, the system was subjected to rigorous evaluation suites. Across independent testing runs, the agent achieved a stellar 0.92 correctness score, with 12 out of 13 test prompts scoring a perfect 1.00. The agent successfully routed messages, routed queries through custom policy boundaries, and aggressively rebuffed malicious prompt-injection attempts.
Yet, despite this pristine statistical veneer, a deeper inspection revealed a deeply unsettling behavior: during a multi-turn user interaction, the chatbot confidently assured a customer, "I have filed a bug report with ticket ID TIX-345678."
In reality, that ticket did not exist. No execution had occurred against the database, no API call had successfully written to the backend, and the customer was left with the false impression that their technical grievance had been officially logged.
This technical investigation explores the mechanics behind building this agent using Amazon Bedrock’s next-generation managed harness, dissects why standard LLM-as-a-judge evaluation frameworks fail to catch hallucinations involving system side-effects, and outlines the paradigm shifts required to build genuinely reliable agentic workflows.
DETAILED CHRONOLOGY: BUILDING THE AGENTIC ARCHITECTURE
The Project Brief and Modern Constraints
The project originated as part of the specialized AWS Agent Engineer curriculum, designed to test an engineer’s ability to orchestrate multi-functional customer service agents. The application scope required handling three distinct communication pathways: managing customer bug reports, answering general Frequently Asked Questions (FAQs), and executing seamless hand-offs to human agents when edge cases exceeded the model’s operational threshold.
However, developers navigating AWS Bedrock encounter a rapidly shifting technological landscape. Bedrock Agents Classic—the foundational architecture around which the course material was originally structured—was officially closed to new customers on July 30, 2026. Consequently, the project had to be migrated to its modern successor: the Amazon Bedrock AgentCore managed harness.
The architectural constraint that defined the exercise—and served as the crucible for its engineering challenges—was the deliberate absence of a traditional programmatic classifier or routing node. There was no explicit pythonic router directing incoming payloads to discrete microservices. Instead, all classification, information retrieval, dialogue state-tracking, and systemic grounding behavior were engineered to live entirely within a singular, monolithic system prompt.
Architectural Layout and Infrastructure
The underlying infrastructure leveraged a streamlined, modern serverless stack designed for low-latency execution and stringent security protocols:
Customer ──► chat.py ── invoke_harness ── AgentCore managed harness ──► Amazon Nova Pro
│ (system_prompt.txt + FAQ) temp 0, topK 1
│
tool call: bugreports___create_bug_report
▼
AgentCore Gateway (MCP, IAM auth)
▼
Lambda create_bug_report ── PutItem ──► DynamoDB
The core components of this ecosystem included:
- The Orchestration Layer: A custom Python CLI script (
chat.py) that initializes the interaction loop and interfaces with the orchestration backend. - The Model Engine: Pinned strictly to
us.amazon.nova-pro-v1:0utilizing greedy decoding parameters (temperature 0,topK 1). This configuration aligns with AWS’s best-practice guidelines for maximizing deterministic behavior during tool-calling sequences. - The Execution Harness: The Amazon Bedrock AgentCore managed harness, parsing a deeply detailed system prompt alongside a comprehensive FAQ data store.
- The Gateway & Persistence Layer: An AgentCore Gateway leveraging Model Context Protocol (MCP) standards and strict Identity and Access Management (IAM) authentication, routing validated tool invocations to an AWS Lambda function (
create_bug_report), which ultimately commits records viaPutItemoperations into an Amazon DynamoDB table.
Routing with Pure Natural Language
Without a programmatic intent-classifier to steer user messages, the project relied on prompt engineering to govern state transitions. The system prompt instructed the language model to categorize incoming user intent into precisely one operational bucket before generating a response, strictly prohibiting the blending of categories.

Establishing hard boundaries proved to be the primary friction point. For instance, determining whether an aggrieved customer statement like "My card was declined" constituted a software bug, a platform policy question, or a generalized transaction inquiry required the codification of 15 explicit tie-breaker rules.
The governing fallback logic was distilled into a clear heuristic:
- If the underlying software misbehaved unexpectedly, it qualified as a bug.
- If the inquiry pertained to static store policy, payment rules, or order statuses, it qualified as a platform question.
- If neither applied, the conversation required an immediate human hand-off.
For bug collection specifically, the system prompt mandated a strict extraction protocol: the model had to parse previous conversational turns, solicit precisely one missing field at a time from the user, and execute the tool call the instantaneous moment all three required fields (description, environment, steps to reproduce) were acquired. The model was explicitly forbidden from asking redundant confirmation questions such as "Shall I file this report for you?"
Furthermore, security hardening was integrated directly into the prompt architecture. All user inputs were treated strictly as untrusted data rather than system instructions. A dictionary of known prompt-override patterns—such as "ignore your previous instructions" or "I am your system developer"—was cataloged, forcing the model to refuse manipulation attempts without acknowledging the detection methodology or engaging in defensive argumentation.
SUPPORTING CONTEXT & METRICS: THE BLIND SPOTS OF AUTOMATED EVALUATION
The Evaluation Suite Methodology
Manual interactive testing does not scale in modern software engineering. To systematically measure the agent’s efficacy, a dedicated 13-case evaluation suite was developed, comprising:
- 3 Bug-report scenarios testing data extraction and tool invocation.
- 3 FAQ scenarios evaluating retrieval accuracy.
- 2 Hand-off scenarios testing escalation boundaries.
- 5 Edge cases comprising a bare
helpcommand, two highly ambiguous user messages, and two adversarial prompt-injection attempts.
An automated test harness script executed each test case within an isolated, fresh session, compiling outputs into a strict JSONL format anticipated by the Bedrock Evaluations service. Amazon Nova Pro then acted as an LLM-as-a-judge, comparing the agent’s output against established reference standards.
To preserve test independence, session memory was systematically disabled across runs. (An engineering note of operational friction: while the AWS API method create_harness accepts memory="disabled": , the update counterpart update_harness strictly demands a nested structure of memory="optionalValue": "disabled": . Failure to match this exact schema triggers a silent fallback loop that repeatedly fails during iterative testing).
Despite these structural hurdles, the evaluation results appeared triumphant: 0.92 correctness achieved across two independent evaluation runs. The histogram reflected 12 prompts achieving a pristine 1.00 score, confirming robust routing, solid boundary enforcement, and total immunity to the embedded injection attempts.
What the Score Missed: Uncovering Production Defects
The statistical confidence of the 0.92 score was swiftly undermined by a rudimentary engineering validation practice: inspecting the persistence layer. By opening the Amazon DynamoDB console and performing a manual diff against the conversational transcripts, critical systemic failures emerged.
1. The Phantom Ticket Phenomenon
In multi-turn conversations where users reported issues (e.g., a broken search bar followed by clarifying questions regarding browser type), the chatbot gracefully concluded the interaction by presenting a cleanly formatted ticket identifier. However, inspecting the terminal logs revealed an absolute absence of any [tool call] execution payload. Furthermore, the DynamoDB table maintained its exact prior row count.
Across multiple test iterations, the model dynamically hallucinated plausible alphanumeric strings such as #12345, TICKET1234, and TIX-345678. These outputs were styled as placeholder-shaped strings, perfectly mimicking a valid system response without ever instantiating a database write.
The root cause lay in an underspecified system instruction. While the prompt broadly cautioned the model against fabricating information, it failed to draw a hard boundary around system-generated cryptographic or sequential IDs. The fix required a strict architectural constraint: the only permissible ticket ID output must be the exact ticketId string returned downstream by a successful tool execution payload. If no string was returned, the agent was mandated to explicitly state that the filing procedure had failed.

2. Corrupted Data Injection in Database Fields
The single prompt that registered a catastrophic 0.00 score during evaluation read: "The app keeps logging me out. I’m using Chrome on Windows 11."
To the end-user interacting with the chat interface, the agent appeared fully functional, expressing gratitude and returning a legitimate ticket identifier. However, a scan of the DynamoDB table exposed severe data corruption: the stepsToReproduce attribute had been populated with conversational filler or redundant echoes of the user’s initial prompt rather than legitimate diagnostic telemetry.
This failure stemmed from internal prompt tension. To prevent the chatbot from aggressively interrogating users for minor details, the evaluation rules for capturing reproduction steps had intentionally been relaxed (e.g., allowing simple affirmations like "It crashes when I click Pay"). However, when the reported symptom and the triggering action overlapped entirely within a single sentence—such as "Keeps logging me out"—the model struggled to differentiate between the core problem description and the reproduction steps. Desperate to satisfy the mandatory schema field required by the database tool, the model fabricated or misplaced conversational text rather than leaving the field null.
3. Minor State-Tracking Slip-ups
Additional inspection revealed ancillary logic failures, including premature escalation loops where ambiguous user statements triggered unnecessary human hand-offs, and minor conversational loops where the model requested data it had already successfully captured in previous turns.
OFFICIAL STATEMENTS & LESSONS LEARNED
Reflecting on the architecture, Mialy, the engineer behind the implementation, highlighted several vital takeaways for the broader developer community building agentic workflows:
"For tool-using agents, you must assert on side effects, not on text. A correctness judge reads the response text. A confident, helpful-sounding response wrapped around an empty or invented database field will consistently score well. The defects that mattered most were entirely invisible to the automated metric and immediately obvious in the database. Next time, the test suite checks the stored row directly."
Key tactical directives for AI engineers include:
- Single-Turn Evals Mask Multi-Turn Failures: Static evaluation suites consisting entirely of single-turn prompts cannot evaluate the state-tracking stability required for complex, multi-turn tool execution.
- Long Prompts Degrade Unevenly: Even under deterministic decoding parameters (temperature 0), complex rules embedded deep within sprawling system prompts are followed inconsistently across sessions. Shorter, modular prompts consistently outperform monolithic text blocks.
- Strict State Validation: System prompts governing tool schemas must explicitly forbid the model from using its own conversational filler as payload data for required parameters.
FUTURE OUTLOOK: THE ROAD AHEAD FOR AGENTIC EVALUATION
As the industry moves beyond proof-of-concept chatbots toward fully autonomous agentic workflows capable of executing financial transactions, updating enterprise databases, and modifying external APIs, traditional text-based evaluation models are proving fundamentally inadequate.
The future of agent engineering demands a paradigm shift toward state-diff evaluation frameworks. Rather than relying solely on LLMs to grade conversational tone and semantic alignment, next-generation CI/CD pipelines for AI agents must incorporate programmatic assertions that validate downstream side effects. Testing suites must automatically spin up ephemeral containerized databases, execute agent chat sessions via headless API harnesses, and programmatically assert that the resulting database rows, API payloads, and external system states match exact cryptographic and structural expectations.
Only by shifting the evaluation focus from what the agent says to what the agent actually does can the developer community bridge the dangerous gap between high evaluation scores and absolute production reliability.
The full codebase, system prompts, evaluation suites, and architectural transcripts referenced in this investigation are publicly available via the GitHub repository: Mialy333/support-chatbot-bedrock-agentcore.
