GitHub introduces ReviewBench for AI code review evaluation
GitHub has released ReviewBench, an open benchmark designed to evaluate AI code review tools. The benchmark is modeled on data from over 100 million real-world pull requests and allows for comparisons based on language, repository size, and code change scope. ReviewBench features a validated set of 219 pull requests across 19 languages and uses multiple sources, including human and automated reviewers, to assess the quality of AI code review systems.
- Modelled on 100M+ GitHub pull requests
- Includes 219 public pull requests, 19 languages
- Severity and category labeling for each finding
- Uses human, LLM, and static analysis reviewers
Sources covering this
More in AI
OpenAI trials visual ads in ChatGPT's image generator
OpenAI will start testing visual ads in ChatGPT this month, showing paid product images while users generate AI images.
Reflection launches Beam, a large open AI model with cheaper compute
Reflection, a well-funded US AI startup, released Beam, a massive open-weight AI model it claims rivals China’s best open models in…
OpenAI speeds up GPT-6 Astra and Sol models
OpenAI made its GPT-6 Astra and Sol models about 50% faster, now generating around 50 tokens per second instead of 30.
Professor swaps coding tests for AI-guided interviews
A university professor has replaced traditional coding exams with interviews that test how students use AI tools like Gemini in Google…