#1754361: Webinar - Benchmark Scores Are a False Flag
| Description: |
Nearly every claim about a model's cybersecurity capability rests on a solve rate. The number counts how many challenges an agent finished, but records nothing about what the agent did to finish them. That gap hides what matters most. A solve rate can't tell you whether a model didn't know the exploit or knew it and failed to execute. It can't tell you whether the agent exploited the vulnerability the challenge was built around or found an exposed credential and took the flag with that instead. And it says nothing about whether the agent stayed inside the target. Tarun Koyalwar, AI Security Researcher at ProjectDiscovery, ran open and closed models against 54 black-box web targets. He gave them no source code, no hints, and no methodology. Then he read every run by hand instead of scoring it. This session covers what he found and the harder question underneath it. The industry already knows these benchmarks are saturated. The problem starts earlier, in how we build them. If you write the test from an answer key and work backward, what are you measuring? What you’ll learn: Why a solve rate is a false flag, and the four dimensions worth measuring in its place How answer-key benchmark design produces challenges that hand over the flag with no exploitation at all What the run-by-run numbers show, including the loaded harness that underperformed an empty one and what repeat runs do to a reported score Why these benchmarks reward an agent for breaking out of scope, and how that connects to the Hugging Face incident Speakers: Davis Franklin, Head of Technical Solutions, ProjectDiscovery Davis Franklin is the Founding Solutions Engineer at ProjectDiscovery. He's based in Boston, MA with over 8 years of experience in cybersecurity and all things AI. His path into cybersecurity ran through physics and software development, giving him a different lens than most. He thinks in first principles, builds what he needs, and knows his way around an attack surface as well as a codebase. These days, he's deep in the weeds with Neo, ProjectDiscovery's AI-powered security testing platform, working hands-on with practitioners to push the limits of what automated security execution can do. Tarun Koyalwar, AI Researcher, ProjectDiscovery Tarun Koyalwar is an AI Researcher for Offensive Security at ProjectDiscovery, where he works on the tracing, evals, and benchmarking behind Neo, PD's offensive security agent. He's self-taught, starting with CTFs, TryHackMe, and web security labs before reporting over 50 vulnerabilities in his first year of bug bounty in 2021. He joined ProjectDiscovery to build and improve offensive tooling, and is a core contributor and maintainer of Nuclei as well as the author of Alterx and Vulnx. He has presented at DEF CON's Bug Bounty Village, Black Hat Asia Arsenal, and BSides Ahmedabad. At BSides Las Vegas 2026. Moderator: Michael Krieger, Moderator, Dark Reading Michael has been involved in various disciplines in the information technology field for more than 25 years. His background and experience in product development, marketing and sales give him a unique perspective on how and why companies succeed in the technology marketplace. He has served as a senior VP for marketing at FutureLink, a large ASP based in Southern California. Before that he was VP of marketing for Hitachi Data Systems’ server product line, and also was a senior product manager for AST Research's server and data communications product lines. |
|---|---|
| More info: | https://dr-resources.darkreading.com/free/w_prpa271/ |
| Date added | Sept. 18, 2026, 6:04 p.m. |
|---|---|
| Source | DarkReading |
| Subjects | |
| Venue | Oct. 8, 2026, midnight - Oct. 8, 2026, midnight |
