JetBrains Research
Research is crucial for progress and innovation, which is why at JetBrains we are passionate about both scientific and market research
Research
AI coding agents can now write thousands of lines of production code across a dozen files in the time it takes you to make coffee. The bottleneck is no longer code generation; it’s review, and nobody’s quite sure how to do it well yet.
Our Human-AI eXperience team, in collaboration with researchers from Lund University, examine how a tool for reviewing AI-generated code should look like and be integrated into the development process. The study is written up in a new paper and will be presented at the Empirical Software Engineering International Week 2026 (ESEIW) in October.
The main message of the study is that the core problem with reviewing AI-generated code is trust calibration. Until the tools catch up, developers will continue to fly blind.
Our proposes a conceptual framework for building AI-ready code review tools that reveal risk and confidence signals at the actual granularity where developers put their attention. This proposal is based on a participatory design study with active feedback and collaboration from 17 practitioners, plus a follow-up survey with 43 software professionals. In this blog post, we present the conceptual framework and then discuss how it applies to developer tools for designing with this framework.
Why your usual review instincts fail with AI-generated code
Traditionally, when you review a colleague’s code, you have a lot of invisible scaffolding working in your favor. You know roughly how senior they are, which parts of the codebase they’re confident in, and which parts they tend to rush. You can ask them questions. When something looks odd, you can ping them on Slack and get a rationale in thirty seconds. However, none of that exists when the author is an LLM.
On top of that, a language model presents every line of generated code with the same apparent confidence, regardless of how uncertain it actually was when it produced that line. There’s no signal telling you that the authentication logic was straightforward and the database migration was a stretch. It all looks the same, and it all reads the same.
The rational response to this is to read every line, because any line could be the one that’s wrong. But, that’s a line-by-line audit of potentially thousands of lines of code, and it scales terribly as AI agents get more capable and start generating larger and larger change sets.
In our we argue that review of AI-generated code needs to be reframed to keep up with the changing times. We frame this as the diff-view paradigm failing to scale. A traditional diff viewer assumes that the reviewer’s job is to understand what changed. But when the author is an LLM that can generate large changes quickly and present heterogeneous output with homogeneous confidence, understanding what changed is the easy part.
Knowing whether to trust it – and where – is the hard part. As AI agents take on increasingly large and complex tasks, and as the volume of LLM-generated code in professional codebases continues to grow, this reframe is going to matter. The diff viewer was the right tool for reviewing what your colleague wrote. It may not be the right tool for reviewing what your agent wrote.
Another way to say it is that reviewing LLM-generated multi-file changes is not a diffing problem, but that it is a trust-calibration problem. We define trust calibration as the capacity to allocate review effort proportionate to segment-level risk when the author can’t be interrogated about their confidence or reasoning. It’s the skill of knowing how to differentiate where to look hard and where you can afford to skim – a skill that human-authored code review supports implicitly through social and contextual cues, and that AI-generated code review currently provides almost no support for at all.
Three-level review workflow: From high-level to selective details
There’s a well-known principle in information visualization from Shneiderman, where the overview comes first, then zoom and filter, then comes the details on demand. This is shown in the image below.
Our three-level review workflow maps onto it almost exactly, following what expert developers say about how they read unfamiliar code. This is sketched in the image below.
Namely, first the reviewer forms high-level hypotheses, then they selectively drill down to validate these hypotheses. In contrast, a more traditional line-by-line diff would force the developer to skip the high-level reviewing, which feels more cognitively expensive when it comes to large amounts of AI-generated changes. That is, when a task contains many interacting elements, a poorly organized presentation consumes processing capacity that could otherwise go toward understanding and judgment.
We developed our proposed workflow structure following a series of four workshops. In addition to a workflow structure, we gleaned recurring design constructs from participants’ inputs, which support the workflow structure. More details on our proposal, and how we brainstormed with developers about what they need in a workflow, can be found in our paper.
Tools: What’s out there now and how our proposal can help tool-builders support developers
None of these ideas appeared in a vacuum, and we discuss this in the paper. Several of the constructs have partial counterparts in tools that already exist: CodeRabbit offers prose walkthrough summaries of pull requests. Claude Code launches multiple reviewer agents that tag findings by severity level. Graphite‘s stacked pull-request model splits large changes into independently reviewable units – the closest existing analogue to our chunk idea, even though it operates across multiple pull requests, rather than decomposing a single generated proposal. GitHub has noticed the shift too. The platform now puts out guidelines specifically for reviewing AI-generated code. It also reports that Copilot-assisted code review already accounts for more than one-fifth of the platform’s reviews.
All of this suggests that the workflow we are proposing is already relevant to developers, even if no one tool fully incorporates the ideas. Especially not in a way that addresses trust calibration. The high-level beginnings to the workflow – overview before files, files before lines, risk stratification before analytical reading – isn’t something any current tool imposes or even suggests. That’s the actual design gap our framework is proposing to fill.
The implication for tool designers is clear and direct: If you’re building IDE tooling for the era of AI-native development, the question to organize around is not How do we display the diff? It’s At what granularity does the reviewer need to allocate attention, and what signal do we need to surface there?
Our three-level framework gives you a principled scaffold for building these tools. Some guidelines are:
- Overview-level tools should substitute for the interpersonal orientation that a human-authored review gets for free, e.g. the ability to ask the author what they were thinking, to draw on knowledge of their strengths and weaknesses.
- File-level tools should do risk stratification before the reviewer reads a single line — helping them invest effort where it matters.
- Code snippet-level tools should recover the fine-grained analytical work that good code review has always required, but equipped with chunk decomposition and chain-of-thought linkage that human-authored review can’t provide.
Our framework also suggests what to avoid when building. That is, tools that address comprehension without addressing trust calibration may improve efficiency for low-stakes changes while leaving the more consequential failure mode untouched. By this, we mean the systematic misallocation of review effort toward low-risk segments and away from high-risk ones.
In a previous blog post, we discussed how extended reality could be integrated into tools to help developers with AI-generated code review (see the section entitled Sketches of what could be built). Keep an eye on this space for updates on similar research.








