Back to Yaps
ConferenceFeatured

Is Your LLM Lying? A Practical Guide to Building Trustworthy AI

Metro Manila, Philippines30 mins

At AWS Community Day Manila 2025 I gave a 30-minute talk about moving from vibes-based testing to real evaluation. An LLM can give a fluent answer that's confidently wrong, and checking answers by hand stops working once you have thousands of conversations. The talk covered how to test, guard, and monitor an LLM app on AWS, mostly in the Bedrock console. The slides are linked above.

What to measure

I sorted "trustworthy" into three groups. Objective metrics are correctness, robustness, and safety. RAG apps add faithfulness, context relevance, and coverage. Then there's subjective quality: helpfulness and coherence.

Testing the base model

Before adding retrieval, test the model on its own, the way you'd unit test a function. Bedrock gives you two ways. A programmatic eval job has you pick the model and a task type (generation, summarization, Q&A, or classification), then scores accuracy, robustness, and toxicity. A model-as-judge job uses a second, usually stronger model to grade answers on a prompt dataset you build, with metrics for quality and for responsible AI.

Testing RAG

RAG evals need more setup. You need your documents in S3, a Bedrock Knowledge Base with an embedding model and a vector store, test questions with expected answers, and the right IAM permissions. Retrieve-only mode tests the retriever by itself and scores context relevance and coverage. Retrieve-and-generate tests the whole pipeline, including whether answers stay faithful to the sources and cite them precisely.

Guardrails

Evals run before release. Guardrails run on live traffic. I demoed contextual grounding, which blocks an answer that isn't supported by the retrieved context and returns a custom message instead. The demo used a grounding threshold of 0.75 and a relevance threshold of 0.60.

Prompt management

Bedrock Prompt Management keeps every version of a prompt, lets you test variants side by side on real data, and promotes the better one without a code change.

Logging

Turn on invocation logging for every model call. Logs go to S3 for long-term storage and to CloudWatch for real-time monitoring, where dashboards track latency, token use, throttling, and usage. Bad answers you find there go back into your test datasets.

Release gates

I closed with the gates I'd want before shipping a change: baseline correctness of at least 80%, context relevance of at least 95%, RAG faithfulness of at least 90%, and guardrails stepping in on no more than 3% of requests.

The console is where you learn the patterns and set baselines. After that, the same checks move into the SDK so they run in CI/CD, Step Functions, or the CLI, and the test set keeps growing from production logs.

Tags

LLMsEvalsGuardrailsAWS Bedrock