Files
krow_backend/go-api/internal/ratelimit/rules.go
Aravind f2aa3b3ad8
Some checks failed
CI / fixture (push) Has been cancelled
CI / test (push) Has been cancelled
mcp connection
2026-09-22 10:58:02 +05:30

164 lines
7.8 KiB
Go

package ratelimit
import "time"
// The rule set for the MCP and OAuth surface, in one place.
//
// Every number below is a judgement, so each carries the reasoning that
// produced it. They are starting values: the right way to change one is to
// change it here, with the comment updated, rather than to pass a different
// number at a call site.
//
// TWO PRINCIPLES SHAPE ALL OF THEM
//
// 1. Limit the scarce thing, not the request. Registration writes a row for an
// anonymous caller, so it is limited hard. A tool call reads rows the
// caller may already read through the product, so it is limited loosely —
// the cost there is database load, not access.
//
// 2. Key by the narrowest identity available. An IP is a whole office behind
// NAT; a token is one connection. Limiting an authenticated endpoint by IP
// would make one person's loop everyone's outage.
var (
// Registration: 10 per hour per IP.
//
// The tightest limit here, because /oauth/register is the only endpoint
// that WRITES for a caller with no credential at all — RFC 7591 requires
// exactly that. A legitimate client registers once per installation and
// then never again, so ten is already generous by two orders of magnitude;
// it is set there only so a developer retrying a broken integration does
// not lock themselves out.
//
// Keyed by IP because there is nothing else to key by: the caller is
// anonymous by definition at this point.
OAuthRegister = Rule{Name: "oauth.register", Limit: 10, Window: time.Hour}
// Authorization: 20 per hour per IP+user.
//
// A person clicking Approve does it once. Twenty allows for a browser
// reload, a mistyped password, a client retrying a flow, and a developer
// testing — and stops a script walking the authorization endpoint to farm
// consent pages or probe client ids.
//
// IP AND user, not either alone: keying by user only would let one
// attacker burn an innocent person's budget by naming them, and keying by
// IP only would make an office share one person's allowance.
OAuthAuthorize = Rule{Name: "oauth.authorize", Limit: 20, Window: time.Hour}
// Token exchange: 30 per hour per client.
//
// One exchange per authorization, and an authorization is already limited
// above — so this is not the primary defence. It is here to bound
// brute-forcing a code or a verifier: an authorization code lives 60
// seconds and is single-use, and 30 attempts an hour makes guessing one
// hopeless rather than merely improbable.
OAuthToken = Rule{Name: "oauth.token", Limit: 30, Window: time.Hour}
// Refresh: 60 per hour per token family.
//
// An access token lives 15 minutes, so a well-behaved client refreshes
// about 4 times an hour. Sixty leaves room for a client that refreshes
// eagerly, or one running several sessions, while bounding a loop.
//
// Keyed by FAMILY rather than by token, because the token changes on every
// rotation — keying by token would give each rotation a fresh budget,
// which is the same as no budget at all.
OAuthRefresh = Rule{Name: "oauth.refresh", Limit: 60, Window: time.Hour}
// MCP tool calls: 60 a minute, and 1000 an hour, per token.
//
// BOTH, because they stop different things. The minute limit stops a tight
// loop — a model retrying a failing call, or a bug — from becoming a spike.
// The hour limit stops a slow, sustained drain that would sit under the
// minute limit forever: 59 calls a minute is 3,540 an hour, which is a lot
// of queries for one connection.
//
// Sixty a minute is well above interactive use. A person asking questions
// generates a handful of calls per turn, and a model doing several lookups
// for one answer still lands in single figures.
MCPToolCallPerMinute = Rule{Name: "mcp.call.min", Limit: 60, Window: time.Minute}
MCPToolCallPerHour = Rule{Name: "mcp.call.hour", Limit: 1000, Window: time.Hour}
// Per-organisation ceiling: 5000 an hour.
//
// The backstop for the case the per-token limits cannot see: one tenant
// with many connected clients, each individually well-behaved, together
// saturating the database. Set well above the sum of a few active users so
// it is never reached in ordinary use — it exists to bound a runaway, not
// to ration normal work.
MCPPerOrgPerHour = Rule{Name: "mcp.org.hour", Limit: 5000, Window: time.Hour}
)
// A note on what is NOT rate limited here, and why.
//
// CONCURRENT CONNECTIONS. The plan proposed 10 concurrent MCP connections per
// user. That is not implemented, and it is not an oversight: this transport is
// stateless — one POST per message, no session, nothing held open — so there is
// no such thing as a concurrent connection to count. The thing that limit was
// reaching for is request rate, and the two limits above are that, measured
// directly. Implementing a connection counter over a stateless endpoint would
// mean inventing connection state purely so it could be limited.
//
// DISCOVERY. The two .well-known documents are static, cacheable for five
// minutes, and contain public URLs. Limiting them would add a database write to
// the cheapest endpoints on the surface, to protect nothing.
//
// REVOCATION. Deliberately unlimited. Revocation is the thing a person reaches
// for when something has gone wrong, and an attacker gains nothing by calling
// it — the worst they can do is revoke tokens they already hold. Rate limiting
// the emergency brake is the wrong trade.
/*
FAILURE BEHAVIOUR, RULE BY RULE
===============================
The question this section answers: when the database cannot be reached, does a
request get through?
EVERY RULE HERE FAILS CLOSED. Limiter.failOpen defaults to false and nothing in
this service sets it to true. The reasoning is the same for all of them and is
worth stating once rather than per-rule:
- A limiter that cannot count is not limiting. If a database outage lifted
the limits, then the moment the system is least able to absorb load is
exactly the moment its protections switch off — and an attacker who can
cause or wait for a blip gets an unmetered window on the endpoints that
write rows for anonymous callers.
- The cost of failing closed is bounded and visible: MCP returns 429 and
Claude retries. The cost of failing open is unbounded and silent.
- These endpoints are not load-bearing for the product. If the database is
down, /oauth/token cannot mint a token and /mcp cannot read a row anyway;
the limiter refusing first changes the error message, not the outcome.
WHAT IS EXPLICITLY NOT FAIL-OPEN, AND WHY IT MATTERS MOST
oauth.register Writes a row for a caller with no credential. Failing open
here is an unauthenticated write endpoint with no ceiling.
oauth.token Bounds brute-forcing a code or a verifier. Failing open
turns a 60-second, single-use code into one an attacker may
guess at without limit for the duration of the outage.
oauth.refresh Failing open removes the bound on a loop against a
long-lived credential.
THE ONE PLACE FAIL-OPEN WOULD BE DEFENSIBLE
A deployment that would rather serve MCP degraded than refuse it can call
WithFailOpen(true) on the limiter used for the mcp.* rules only — those guard
database load rather than access, and every call behind them is already
authenticated and already authorized by the policy table. That is a deliberate
operational trade, it is one line, and it is deliberately not the default.
It must NOT be applied to the oauth.* rules. Those guard the credential issuance
path, where the thing being limited is an attacker's number of attempts.
OBSERVABILITY
A limiter failure is logged at ERROR by the middleware (httpserver/mcplimit.go)
with the rule name and the decision, never the subject — the subject is a hash
of a credential. A sustained run of those log lines means the limiter is not
limiting, and is worth an alert.
*/