164 lines
7.8 KiB
Go
164 lines
7.8 KiB
Go
package ratelimit
|
|
|
|
import "time"
|
|
|
|
// The rule set for the MCP and OAuth surface, in one place.
|
|
//
|
|
// Every number below is a judgement, so each carries the reasoning that
|
|
// produced it. They are starting values: the right way to change one is to
|
|
// change it here, with the comment updated, rather than to pass a different
|
|
// number at a call site.
|
|
//
|
|
// TWO PRINCIPLES SHAPE ALL OF THEM
|
|
//
|
|
// 1. Limit the scarce thing, not the request. Registration writes a row for an
|
|
// anonymous caller, so it is limited hard. A tool call reads rows the
|
|
// caller may already read through the product, so it is limited loosely —
|
|
// the cost there is database load, not access.
|
|
//
|
|
// 2. Key by the narrowest identity available. An IP is a whole office behind
|
|
// NAT; a token is one connection. Limiting an authenticated endpoint by IP
|
|
// would make one person's loop everyone's outage.
|
|
|
|
var (
|
|
// Registration: 10 per hour per IP.
|
|
//
|
|
// The tightest limit here, because /oauth/register is the only endpoint
|
|
// that WRITES for a caller with no credential at all — RFC 7591 requires
|
|
// exactly that. A legitimate client registers once per installation and
|
|
// then never again, so ten is already generous by two orders of magnitude;
|
|
// it is set there only so a developer retrying a broken integration does
|
|
// not lock themselves out.
|
|
//
|
|
// Keyed by IP because there is nothing else to key by: the caller is
|
|
// anonymous by definition at this point.
|
|
OAuthRegister = Rule{Name: "oauth.register", Limit: 10, Window: time.Hour}
|
|
|
|
// Authorization: 20 per hour per IP+user.
|
|
//
|
|
// A person clicking Approve does it once. Twenty allows for a browser
|
|
// reload, a mistyped password, a client retrying a flow, and a developer
|
|
// testing — and stops a script walking the authorization endpoint to farm
|
|
// consent pages or probe client ids.
|
|
//
|
|
// IP AND user, not either alone: keying by user only would let one
|
|
// attacker burn an innocent person's budget by naming them, and keying by
|
|
// IP only would make an office share one person's allowance.
|
|
OAuthAuthorize = Rule{Name: "oauth.authorize", Limit: 20, Window: time.Hour}
|
|
|
|
// Token exchange: 30 per hour per client.
|
|
//
|
|
// One exchange per authorization, and an authorization is already limited
|
|
// above — so this is not the primary defence. It is here to bound
|
|
// brute-forcing a code or a verifier: an authorization code lives 60
|
|
// seconds and is single-use, and 30 attempts an hour makes guessing one
|
|
// hopeless rather than merely improbable.
|
|
OAuthToken = Rule{Name: "oauth.token", Limit: 30, Window: time.Hour}
|
|
|
|
// Refresh: 60 per hour per token family.
|
|
//
|
|
// An access token lives 15 minutes, so a well-behaved client refreshes
|
|
// about 4 times an hour. Sixty leaves room for a client that refreshes
|
|
// eagerly, or one running several sessions, while bounding a loop.
|
|
//
|
|
// Keyed by FAMILY rather than by token, because the token changes on every
|
|
// rotation — keying by token would give each rotation a fresh budget,
|
|
// which is the same as no budget at all.
|
|
OAuthRefresh = Rule{Name: "oauth.refresh", Limit: 60, Window: time.Hour}
|
|
|
|
// MCP tool calls: 60 a minute, and 1000 an hour, per token.
|
|
//
|
|
// BOTH, because they stop different things. The minute limit stops a tight
|
|
// loop — a model retrying a failing call, or a bug — from becoming a spike.
|
|
// The hour limit stops a slow, sustained drain that would sit under the
|
|
// minute limit forever: 59 calls a minute is 3,540 an hour, which is a lot
|
|
// of queries for one connection.
|
|
//
|
|
// Sixty a minute is well above interactive use. A person asking questions
|
|
// generates a handful of calls per turn, and a model doing several lookups
|
|
// for one answer still lands in single figures.
|
|
MCPToolCallPerMinute = Rule{Name: "mcp.call.min", Limit: 60, Window: time.Minute}
|
|
MCPToolCallPerHour = Rule{Name: "mcp.call.hour", Limit: 1000, Window: time.Hour}
|
|
|
|
// Per-organisation ceiling: 5000 an hour.
|
|
//
|
|
// The backstop for the case the per-token limits cannot see: one tenant
|
|
// with many connected clients, each individually well-behaved, together
|
|
// saturating the database. Set well above the sum of a few active users so
|
|
// it is never reached in ordinary use — it exists to bound a runaway, not
|
|
// to ration normal work.
|
|
MCPPerOrgPerHour = Rule{Name: "mcp.org.hour", Limit: 5000, Window: time.Hour}
|
|
)
|
|
|
|
// A note on what is NOT rate limited here, and why.
|
|
//
|
|
// CONCURRENT CONNECTIONS. The plan proposed 10 concurrent MCP connections per
|
|
// user. That is not implemented, and it is not an oversight: this transport is
|
|
// stateless — one POST per message, no session, nothing held open — so there is
|
|
// no such thing as a concurrent connection to count. The thing that limit was
|
|
// reaching for is request rate, and the two limits above are that, measured
|
|
// directly. Implementing a connection counter over a stateless endpoint would
|
|
// mean inventing connection state purely so it could be limited.
|
|
//
|
|
// DISCOVERY. The two .well-known documents are static, cacheable for five
|
|
// minutes, and contain public URLs. Limiting them would add a database write to
|
|
// the cheapest endpoints on the surface, to protect nothing.
|
|
//
|
|
// REVOCATION. Deliberately unlimited. Revocation is the thing a person reaches
|
|
// for when something has gone wrong, and an attacker gains nothing by calling
|
|
// it — the worst they can do is revoke tokens they already hold. Rate limiting
|
|
// the emergency brake is the wrong trade.
|
|
|
|
/*
|
|
FAILURE BEHAVIOUR, RULE BY RULE
|
|
===============================
|
|
|
|
The question this section answers: when the database cannot be reached, does a
|
|
request get through?
|
|
|
|
EVERY RULE HERE FAILS CLOSED. Limiter.failOpen defaults to false and nothing in
|
|
this service sets it to true. The reasoning is the same for all of them and is
|
|
worth stating once rather than per-rule:
|
|
|
|
- A limiter that cannot count is not limiting. If a database outage lifted
|
|
the limits, then the moment the system is least able to absorb load is
|
|
exactly the moment its protections switch off — and an attacker who can
|
|
cause or wait for a blip gets an unmetered window on the endpoints that
|
|
write rows for anonymous callers.
|
|
|
|
- The cost of failing closed is bounded and visible: MCP returns 429 and
|
|
Claude retries. The cost of failing open is unbounded and silent.
|
|
|
|
- These endpoints are not load-bearing for the product. If the database is
|
|
down, /oauth/token cannot mint a token and /mcp cannot read a row anyway;
|
|
the limiter refusing first changes the error message, not the outcome.
|
|
|
|
WHAT IS EXPLICITLY NOT FAIL-OPEN, AND WHY IT MATTERS MOST
|
|
|
|
oauth.register Writes a row for a caller with no credential. Failing open
|
|
here is an unauthenticated write endpoint with no ceiling.
|
|
oauth.token Bounds brute-forcing a code or a verifier. Failing open
|
|
turns a 60-second, single-use code into one an attacker may
|
|
guess at without limit for the duration of the outage.
|
|
oauth.refresh Failing open removes the bound on a loop against a
|
|
long-lived credential.
|
|
|
|
THE ONE PLACE FAIL-OPEN WOULD BE DEFENSIBLE
|
|
|
|
A deployment that would rather serve MCP degraded than refuse it can call
|
|
WithFailOpen(true) on the limiter used for the mcp.* rules only — those guard
|
|
database load rather than access, and every call behind them is already
|
|
authenticated and already authorized by the policy table. That is a deliberate
|
|
operational trade, it is one line, and it is deliberately not the default.
|
|
|
|
It must NOT be applied to the oauth.* rules. Those guard the credential issuance
|
|
path, where the thing being limited is an attacker's number of attempts.
|
|
|
|
OBSERVABILITY
|
|
|
|
A limiter failure is logged at ERROR by the middleware (httpserver/mcplimit.go)
|
|
with the rule name and the decision, never the subject — the subject is a hash
|
|
of a credential. A sustained run of those log lines means the limiter is not
|
|
limiting, and is worth an alert.
|
|
*/
|