Bot monitoring — free

Your AI reviewer has an invoice. This is the receipt.

Every bot comment on your repos, graded by an independent model — severity, category, cost and overlap. Whether the bot seat earns its keep stops being a feeling and becomes a number you can act on.

01 / The model

Graded by a model with no seat to sell.

Bot vendors grade their own homework — five of them currently claim #1 on the same public benchmark, citing different versions and metrics. And a public leaderboard, however honest, can’t tell you what happened on your repos with your team. Limn’s classifier is its own: fine-tuned on a multi-year corpus of GitHub bot reviews, running on plain CPU in the hosted service — not the bot’s self-assessment, and not an LLM invoice. Every bot-authored comment gets a severity and up to eight categories, seconds after sync.

It never competes with your bots. Limn doesn’t review code — it measures the reviewers, which is exactly why the measurement can be trusted. The meta-layer above every bot, including anything Limn itself runs.

Honestly, though

Labels are advisory. Severity buckets the top of the scale (major + critical read as “high”), the nit/minor boundary is genuinely fuzzy, and confidence is shown — nothing auto-acts on a label. A grader that hid its uncertainty would be one more bot to distrust.

02 / The mix

What your bots actually said.

The rollup answers the noise question per bot, per sprint: what share was nitpick, style, correctness, security, tests. A bot that’s 61% nits isn’t necessarily a bad bot — but it’s a fact worth knowing before renewal, and a tuning target you can point at.

“~19% were good, 2% were flat-out incorrect, and 79% were nits”
Greptile, measuring its own review bot’s comments — vendor blog, December 2024

On the PR itself, every bot comment wears its grade — take the high-severity flags first, sweep the nits in one pass.

Sprint review · “The bot feels noisy” arrives at the meeting as a chart, not an argument.

limn · graded threads
03 / Value

Cost per useful comment.

Type each bot’s monthly price into its row and the receipt fills itself: volume, severity mix, the share of comments a human actually acted on, and what every acted-on comment cost you. Review bots run $24–30 a seat — whether that’s a bargain or a subsidy for noise is now arithmetic. In a DX survey of 50 engineering budget holders, 86% were uncertain which of their AI tools actually provided benefit; this row is that answer, for review bots.

Priced per workspace, so the platform team’s verdict on a bot doesn’t overwrite the web team’s.

Budget season · The renewal email arrives. The answer is already a number on the board.

limn · value for money
04 / Noise

Where the noise comes from — and where nobody’s listening.

Noise concentrates: in a repo, in a bot, in a category. Limn locates it — including the overlap between bots (two vendors raising the same issue on the same code is paying twice to be told the same thing), the bot-only reviews no human ever handled — the clearest sign a bot needs tuning, not a bigger audience — and your in-house bots, which get counted like any vendor. Lint-wrapper or $30 seat, same receipt.

limn · overlap
limn · bot-only reviews
limn · in-house bots

Tuning day · The mix names the worst offender and its loudest category. You tune one rule, and next sprint’s chart shows whether it worked.

05 / The verdict

Keep, tune, or kill — calmly.

Every bot carries a working verdict, computed from what your team actually did with its output — and it’s patient by design: a comment only counts against a bot after a 36-hour grace window, so “unhandled” means ignored, not merely recent. Disagree? Reclassify any bot in a click; your judgement wins over the detection, per workspace.

Coming: the opt-in, anonymised cross-org benchmark — your bots’ acted-on rate against teams like yours. The receipt gets a reference column.

Themes & per-severity reports live in Pro →

Quarter end · Keep two, tune one, cancel one — decided in a minute, receipt attached.

limn · bot settings

Free. Because you’d doubt a ruler you rented.

The receipt is in the free tier — sign in with GitHub, point it at your repos, and read this month’s number. Grading runs in the hosted service today; bringing the model to the local install is on the roadmap.