Skip to content

Concepts

Authentication and limits

Bearer keys, OAuth on the MCP endpoint, who owns what, and the request and concurrency limits a key runs under.

Two ways in, one identity behind them. A REST client sends a bearer key; an assistant connecting over MCP goes through OAuth. Either way the request arrives as the same caller, spending the same organization’s balance, under the same two limits.

Bearer keys

Keys are created in the console and shown exactly once, at creation. We store only a hash afterwards, so a lost key cannot be recovered and a stolen database cannot be replayed against the API — if you lose one, issue another and revoke the old one. Send it as a bearer token on every request:

Authorization: Bearer sk_live_YOUR_KEY

The console shows the first few characters of each key so you can tell them apart, and revoking one takes effect on the next request. Issue a key per integration rather than sharing one: revoking is the only way to cut something off, and you want that to be a decision about one client.

GET /v1/account/balance·Free.

curl "https://www.shortsintel.com/v1/account/balance" \
  -H "Authorization: Bearer sk_live_YOUR_KEY"

GET /v1/account/balance · Free. The cheapest way to check a key works.

OAuth on the MCP endpoint

The MCP server at https://www.shortsintel.com/api/mcp accepts the same bearer keys, and also speaks OAuth 2.1 so a client can connect without a human pasting a secret into a config file. The client discovers the authorization server from the endpoint, the user signs in, picks the organization the connection acts for, and the client gets a short-lived access token it refreshes on its own.

Access tokens are checked on every request, and so is membership: the organization the user chose at consent is re-read from the database each time, so removing someone from an organization locks their assistant out immediately rather than whenever the token happens to expire. Three codes are specific to this path — invalid_token (401) for any token that fails verification, organization_not_selected (403) when the connection carries no organization to bill, and organization_access_revoked (403) when the user has since left it. The first is fixed by refreshing; the other two need re-authorizing.

The MCP quickstart walks through connecting a client either way.

Who owns what

A key belongs to an organization, not to a person. The organization owns the balance, the Runs, the trackers, the webhook endpoints and the spend log; every key it issues spends that one balance, so a staging key with a loop in it drains the same money production does.

Everything is scoped to that owner and nothing crosses between owners. Asking for another organization’s Run answers not_found (404) rather than a forbidden: you cannot learn that something exists unless it is yours.

GET /v1/runs·Free.

curl "https://www.shortsintel.com/v1/runs?limit=25&status=running" \
  -H "Authorization: Bearer sk_live_YOUR_KEY"

GET /v1/runs · Free. Only ever your own.

The two limits

60 requests per minute per credential — per key, or per user for an OAuth connection, since tokens rotate and a client can hold several. Free reads count toward it, so a tight polling loop is throttled even though it is never charged. Over the line and you get 429 with rate_limit_exceeded:

429 · `rate_limit_exceeded` or `concurrency_limit_exceeded`, with `Retry-After`.

{
  "type": "https://docs.shortsintel.com/errors/rate_limit_exceeded",
  "title": "Rate limit exceeded",
  "status": 429,
  "code": "rate_limit_exceeded",
  "detail": "60 requests per minute per key. Retry after the number of seconds in `Retry-After`.",
  "limit": 60
}

10 Runs in flight at once, counted across the whole organization. Only Runs still running count; finished ones never do. This is the limit that actually protects a balance from a looping agent — the request limit caps how fast calls arrive, this one caps how much paid work can be happening at the same time. Over it you get 429 with concurrency_limit_exceeded instead.

Both answer 429 and both carry Retry-After, but they want different things from you. On rate_limit_exceeded, slow down: sleep the seconds you are given and spread the calls out. On concurrency_limit_exceeded, do not slow down — wait, because the queue only moves when one of your own Runs finishes. Retry the same request after the interval either way; nothing was charged and no Run was created.

The concurrency cap is per organization, so adding keys does not add Run headroom; the request limit is per key (per user over OAuth). If a limit is genuinely too tight for what you are building, the fix is a conversation, not more credentials — the errors guide lists every other code you might meet on the way there.