12 W 39th St
Most teams pick an AI model the same way they pick a coffee machine - the demo looks good or they hear about it from someone, so they assume it works. But in production, evaluation decisions directly impact cost, latency, accuracy, and user trust. Without a structured eval practice, you're flying blind when a model regresses, or a new release claims to be better. This session gives you a framework for evaluating LLMs against *your* workloads and operationalizing evaluation as a continuous practice, not a one-time exercise. **What we'll cover:** * Why standard benchmarks are often misleading for enterprise use cases * Building a golden dataset: how to sample, label, and version evaluation sets from real production traffic * Key metrics beyond accuracy: hallucination rates, precision/recall tradeoffs, cost-per-correct-answer, and latency percentiles * Running evals at scale: open-source tools (Ragas, PromptFoo, LangSmith) and using Amazon Bedrock Model Evaluation * Setting a model evaluation gate: how to enforce a promotion threshold before any model goes to production * Demo: running a side-by-side eval across two models on a real classification task **High level schedule:** * 4:30 - 4:45: Arrivals * 4:45 - 5:45: Presentation & Demo * 5:45 - 6:00: Q&A * 6:00 - 6:30: Networking **Important instructions** * Event starts at 4:45 pm EST, but please allow at least 15 minutes for security to process registration * Please ensure that your meetup profile has your **full name.** Both first and last name are required and we will not be able to register attendees with just abbreviations or incomplete names. * There is an optional networking and Q&A event at the end of the meetup
Free
Thursday, August 27 · 4:30 PM