Is Your LLM Lying? A Practical Guide to Building Trustworthy AI
At AWS Community Day Manila 2025 I gave a 30-minute talk about moving from vibes-based testing to real evaluation. An LLM can give a fluent answer that's confidently wrong, and checking answers by hand stops working once you have thousands of conversations. The talk covered how to test, guard, and monitor an LLM app on AWS, mostly in the Bedrock console. The slides are linked above.
What to measure
I sorted "trustworthy" into three groups. Objective metrics are correctness, robustness, and safety. RAG apps add faithfulness, context relevance, and coverage. Then there's subjective quality: helpfulness and coherence.
Testing the base model
Before adding retrieval, test the model on its own, the way you'd unit test a function. Bedrock gives you two ways. A programmatic eval job has you pick the model and a task type (generation, summarization, Q&A, or classification), then scores accuracy, robustness, and toxicity. A model-as-judge job uses a second, usually stronger model to grade answers on a prompt dataset you build, with metrics for quality and for responsible AI.
Testing RAG
RAG evals need more setup. You need your documents in S3, a Bedrock Knowledge Base with an embedding model and a vector store, test questions with expected answers, and the right IAM permissions. Retrieve-only mode tests the retriever by itself and scores context relevance and coverage. Retrieve-and-generate tests the whole pipeline, including whether answers stay faithful to the sources and cite them precisely.
Guardrails
Evals run before release. Guardrails run on live traffic. I demoed contextual grounding, which blocks an answer that isn't supported by the retrieved context and returns a custom message instead. The demo used a grounding threshold of 0.75 and a relevance threshold of 0.60.
Prompt management
Bedrock Prompt Management keeps every version of a prompt, lets you test variants side by side on real data, and promotes the better one without a code change.
Logging
Turn on invocation logging for every model call. Logs go to S3 for long-term storage and to CloudWatch for real-time monitoring, where dashboards track latency, token use, throttling, and usage. Bad answers you find there go back into your test datasets.
Release gates
I closed with the gates I'd want before shipping a change: baseline correctness of at least 80%, context relevance of at least 95%, RAG faithfulness of at least 90%, and guardrails stepping in on no more than 3% of requests.
The console is where you learn the patterns and set baselines. After that, the same checks move into the SDK so they run in CI/CD, Step Functions, or the CLI, and the test set keeps growing from production logs.