After setting --max-findings 5, how does the system decide which 5 get kept and which get dropped?
Public documentation doesn't detail the specific ranking algorithm, but based on this parameter's design purpose (controlling the cap on findings presented to the reviewer), it's reasonable to infer that the system internally completes full issue detection first, then returns only the top few according to some severity or importance ranking logic — meaning "finding fewer issues" and "finding everything but only reporting the most important" are two different things, and --max-findings does the latter.
If you find a review with a cap of 5 missed an issue you consider important, the practical move is to run it again with all to get the complete list for comparison, to confirm whether the ranking logic diverges from your own judgment, or whether that issue simply wasn't detected in the first place.
If different people on a team set different --max-findings numbers for the same review, could that lead to inconsistent perceived review quality across the team?
In theory, yes, in a specific sense. Assuming the system's internal ranking logic is fixed and doesn't change which issues get judged "most important" based on the number set, someone using a smaller number simply sees a subset of the full list, not a lower-quality review — but in practice, if some people on the team habitually set 3 and others set 20, discussions about "did this review catch issue X" can create confusion from differing list lengths, easily mistaken for inconsistent review quality when it's actually just a difference in presentation scope.
The better approach is for the team to standardize internally — setting a shared default per context (first pass vs. final review) rather than leaving it to individual habit — so that at minimum, everyone sees the same standard within the same context.
Does --max-findings default produce the same result as running /code-review without this parameter at all?
Based on the parameter's design logic, this should be the case — the purpose of the default option is more about letting you explicitly write "use the system default" as an intent in a script or config file, rather than omitting the parameter and letting behavior implicitly fall to the default. For interactive manual command entry, the two should theoretically have no difference in effect; but if you're wrapping /code-review into an automation script, explicitly writing --max-findings default makes it clearer to anyone reading that script later that this is a deliberate choice of default behavior, not a forgotten parameter.
For a team just starting to integrate /code-review into their workflow, what number should --max-findings start at?
There's no universal number, but a reasonable starting point is to run a full review with all a few times first, and actually observe how many findings a typical PR tends to produce and roughly what proportion of them actually get acted on. If you notice a medium-sized PR regularly produces 30-plus findings, but the team in practice only ends up addressing the top 5 to 8 of them, setting your everyday default in that range will more closely match the team's actual review behavior pattern, rather than guessing a number out of thin air.
This calibration process itself is worth redoing periodically — as codebase quality improves over time or the team's review standards shift, a reasonable default may need to change too; it's not a one-time, set-and-forget decision.
When using Claude Code's /code-review command, a common practical frustration is this: running a review on a commit or PR with a large diff can return dozens of findings all at once, everything from serious logic errors to trivial naming suggestions mixed together, leaving the reviewer to do extra work just filtering which of these suggestions are actually worth acting on. A recent release adds a --max-findings parameter that lets you directly control the cap on how many findings a given review returns.
/code-review --max-findings <n>|all|default supports three forms: give a specific number (say, --max-findings 5) to cap this review at returning at most 5 findings; pass all to remove the cap and return every issue found; or pass default to fall back to the system's original default behavior. This parameter lets you switch directly between "broad scope, want a thorough check" and "narrow scope, just want a quick look at the most critical issues" with a single flag, rather than manually filtering a long list yourself every time.
A code review tool's output quality isn't just about whether it catches issues — it's also about whether the presentation creates cognitive burden for the reviewer. Dumping too many findings at once, even if each one is individually reasonable, dilutes the reviewer's attention away from the genuinely important issues — a risk similar to the well-known phenomenon in human code review where too many comments makes people more inclined to just accept everything without careful thought. Capping the number of findings effectively forces the system to rank issues by severity and return only the handful most worth prioritizing, which is particularly useful for scenarios where review needs to happen within limited time.
If you're in a quick first-pass PR review stage, aiming to catch merge-blocking critical issues (security vulnerabilities, logic errors), setting a smaller number (say, 3 to 5) fits better — this forces the system to report only what it considers the most severe items, letting you handle those first before deciding whether a more thorough review is needed. If you're doing a one-off, comprehensive code quality audit (say, a health check before a major refactor), using all to get every finding back is more appropriate, because in that scenario what you actually need is the complete issue list, not a filtered subset.
If your team has already integrated /code-review into your routine PR workflow, --max-findings is worth building into your team's standard operating procedure, rather than leaving it to individual judgment whether to scroll through the full list each time — the more sensible approach is to set different default values for different review contexts (a quick first pass versus a final pre-merge review) and document them in your team's code review guidelines, so review rigor scales with review urgency instead of defaulting to the same "list everything, filter it yourself" approach every time and adding unnecessary review time cost.