Why grr exists
The short version: my inbox hit 9,193 unread, so I made it queryable. The long version is a set of engineering bets about generated surfaces, honest transports, and a repository that ships itself — and two release bugs that took a week each to understand.
There is no mail client built for a queue that happens to be email.
The inbox was never the problem. Querying it was.
My inbox had 9,193 unread emails. I sit on a strata committee, which means hundreds of invoices a year, and most of them arrive by email. When I finally looked at the pile, the texture was specific: newsletters I meant to get to, LinkedIn notifications about profile views, security alerts, and — the one that stung — four of the first twelve unread were my own project's CI robots reporting that a workflow had finished.
None of that is a mail problem. It is a queue problem, and the tools built for queues — grep, jq, scripts — are not built for Gmail. So I did the thing I know how to do: I made the inbox reachable from a terminal.
The first commit was called gmail-opencode-rust: a Rust rewrite of a TypeScript Gmail
tool, born less from a product need than from transport curiosity. I wanted to work with QUIC,
HTTP/3 and streaming, and Gmail was a real API with real rate limits. The first design decision
that survived had nothing to do with mail — do not throttle on the client. Send the request, and
honour the server's Retry-After when it pushes back. Client-side rate limiting is
guesswork about someone else's capacity; the server already knows.
That experiment grew a wider surface than I expected.
Generate the tree. Never hand-write it.
401 methods across 14 APIs cannot be hand-maintained, so they are not.
A hand-written command tree can never keep up with Google Workspace. There are
401 methods across 14 APIs,
and Google adds more on its own schedule. So grr does not hand-write one. Google publishes a
machine-readable Discovery document per API; the repository commits a distilled index — a few
hundred KiB across the 14 services — and
scripts/generate-commands.ts compiles the entire command tree from it:
401 leaves and 994 typed flags, built with clap's builder API rather than
derive, because the tree is composed at runtime and derive cannot express that.
The naming rule is the whole interface: a method id, dots turned into spaces, is the command.
gmail.users.messages.list -> grr gmail users messages list
sheets.spreadsheets.values.get -> grr sheets spreadsheets values get
Parameter names become kebab-case flags (userId → --user-id), and a
parameter that would collide with a global flag surfaces prefixed (format →
--param-format). Because the index is compiled in rather than fetched, grr schema and grr api list answer with no network and no credential —
which means an agent can ask what Sheets can do before it has a login, the order you want
to fail in.
There is one escape hatch, not two implementations. grr api call <id> reaches
any method by id through the same engine as the generated leaf, producing byte-identical
requests. That kept testing honest: a behaviour verified on the typed path is verified on the flat
path too, because they are the same call.
Remove the step where people quit.
A CLI usually dies before its first command. The release build deletes that failure mode.
The most common way a command-line tool dies is at the first run: install, then read a page about creating a Cloud project, enabling a dozen APIs, minting an OAuth client, and pasting two secrets into a config file. Every one of those steps is a place to walk away.
Release builds remove the step. build.rs reads the OAuth client id and secret from
the build environment — GitHub Actions repository secrets for the official binaries — and compiles
them in. grr auth login opens a browser, you consent, and every service works. Google
treats installed-app client secrets as non-confidential, and PKCE is always on, so the flow does
not depend on the secret staying secret. Source builds still bring their own client:
grr auth setup writes one to ~/.grr/config.toml, so contributors and
self-hosters are not locked out — they just opt in.
Report the protocol you actually negotiated.
HTTP/3 is always compiled in with a silent HTTP/2 fallback. The tool tells you which one you got.
HTTP/3 is compiled into every build (reqwest's unstable http3 path, rustls plus quinn), with a
prior-knowledge probe at startup and a silent fallback to HTTP/2. The honest part is that grr transport reports what actually happened, not what the README hopes for:
negotiated_protocol: HTTP_3
http3_requested: true
http3_effective: true
fell_back: false
If QUIC is blocked on your network you get fell_back: true and HTTP/2, and nothing is
wrong — a harness should not retry or treat it as a fault. The point is that the transport claim
is checkable in one command instead of being taken on faith.
The repository maintains itself.
The part I am most pleased with is not a feature. It is that the repo runs itself and ships on every push.
A daily workflow refetches Google's Discovery documents and regenerates the index along with everything derived from it — the command tree, the per-service agent skills, and the site's coverage table — then opens a pull request when anything differs. Another regenerates the changelog from the commit history on every push. Another re-records the terminal demo on CLI-surface pushes, so the demo in the README cannot drift from the binary. Stats and the benchmark refresh nightly.
As of this release, every push to main ships: the version-bump pull request is opened and
auto-merged, the tag is pushed, five platform binaries are built, and the crate publishes to
crates.io through trusted publishing. The weekly run is a backstop, not the cadence — after a
release the next push's recommendation is none, so the loop converges on its own.
This is a bet that the generated surface stays reviewable. A daily pull request that touches the index, the tree, the skills and the coverage table is a real diff to read. If that diff stopped being readable, the whole approach would be worse than hand-writing commands, and I would rather know that than pretend otherwise.
docs.rs failed all five builds with a parse error.
Two layers, both the same shape: configuration that works here and does not travel with the crate.
docs.rs failed all five grr-cli builds with invalid Cargo.toml syntax
. Not a compile error
— a parse error, at the pre-build step, which is the least debuggable kind.
Root cause: trim-paths and panic-immediate-abort are unstable Cargo
profile keys, gated locally by .cargo/config.toml — a file that never travels with a
published crate. On this machine the manifest parses; uploaded to crates.io and run in a bare
environment, plain cargo metadata chokes. That is exactly what docs.rs runs.
The fix is a manifest-level opt-in, cargo-features = ["trim-paths",
"panic-immediate-abort"], the one gate that ships inside the crate. It changes
nothing for consumers, because the crate has always required nightly. Then came layer two: with
the manifest parsed, reqwest's http3 feature refused to compile there, needing --cfg reqwest_unstable — which again came from the config file that does not travel. [package.metadata.docs.rs] now injects the three cfgs through cargo-args,
which reaches dependencies as well as the crate. I verified it by building reqwest's http3 both
ways: fails bare, compiles injected.
The lesson is not read the docs
. It is that works on my machine
has a precise
meaning for published artifacts: the only environment that matters is the one you do not control,
and the config file you rely on is the one that is not in the tarball.
winget validation looped for a week on a file name.
The manifest promised one executable; the archive shipped another. Nothing anyone typed was wrong — that was the problem.
The winget submission sat on issue with installing the application correctly
for a week.
Validation kept failing without naming a file.
Root cause: the Windows release zip contained windows-x86_64.exe — the binary named
after its build target — while the manifest promised grr.exe. Validation could never
find the nested installer, so it looped. The same misshapen name forced every human extracting a
tarball to rename by hand.
The fix is in two places. Archives now stage the binary as grr / grr.exe
inside a per-target directory. And the manifest is no longer written by hand: scripts/generate-winget.ts downloads the released artifact, reads the zip's real
member list, takes the checksum from SHA256SUMS, and derives every field from what
actually shipped. Run against v0.7.0's live release it correctly derives windows-x86_64.exe — both the proof of the root cause and the proof that the
derivation is honest. The manifest can no longer name a file the archive does not contain.
The listing still has to clear Microsoft's validation, and it has not yet. But the class of bug is closed, because the manifest is now a function of the artifact instead of a second copy of it.
Publish the losses, then name what is missing.
A benchmark you only win is marketing. The compare page carries the rows grr loses too.
The site runs a nightly comparison against gog and gws and publishes
the rows where grr loses. gog covers 27 services, including domain-locked enterprise
APIs that a consumer account cannot authorise and that grr deliberately leaves out. gws rebuilds its surface from Google's live Discovery Service at runtime, so it has a
new method before any release, and it ships 100+ agent skills. Those facts sit on the compare page next to the numbers grr wins. As of the current
nightly snapshot, the startup median is about 4 ms and the grr schema dump about 18
ms, with peak memory measured on each tool's closest-equivalent workload rather than a synthetic
one — and a daily job that will change those numbers, which is why this essay links the live page
instead of printing a fixed table.
What is still missing, honestly:
- No client-side throttling, by design. The tool honours
Retry-After, but a script that ignores it can lean on Google's quota harder than it should. That is the deliberate trade: the server knows its capacity, and grr does not second-guess it. - The scope set is fixed at login. Tasks needs a scope this build does not request, so a Tasks call can come back 403 with an error naming the missing scope. It is a real gap, not a rounding error.
- winget is not published yet. The root cause is fixed and the submission is derived, but Microsoft validation has the final say.
- A source build needs Rust nightly, because HTTP/3 sits behind an unstable cfg. That is a real cost and it is not going away while the transport is where it is.
What I want next is boring: close the winget loop, keep the daily pull request easy to review, and keep the demo honest. The bet that matters is unchanged — the command tree is generated from an index committed to the repository, so it works offline, it is reviewable in a diff, and it cannot quietly drift from Google's own description of its APIs. Everything else is maintenance.
Install is one line. The release binary is zero-config:
brew tap debanjanbasu/tap
brew trust debanjanbasu/tap
brew install grr The source, the releases and this site are all linked from the repository ↗. If any claim above looks wrong, the diff, the workflow, or the benchmark that produced it is public.
Where the details live.
This was the story. The reference pages carry the specifics, and the benchmark is the one that stays current.
The generated tree
How 401 methods across 14 APIs become one naming rule, with the typed-flag rules and per-service examples.
The Discovery index
The committed index behind the tree: grr api list, describe, call, refresh, and the per-method scopes.
The benchmark
The nightly comparison against gog and gws, including the rows where grr loses. Numbers change; the page is the source.