Skip to content
Per-key API abuse detection

A stolen key bills you like a customer

It authenticates, it pays, and nothing in your stack finds it strange. Cerberus watches key-level usage metadata and tells you when one key starts being used from many places at once — the shape a credential makes once it has been shared. One signal, deliberately narrow, with its limits published.

Twelve fields · never your keys · never your prompts

What Cerberus receives

Twelve fields, nothing else. The ingest endpoint rejects any unknown field by name, so this list is enforced in code rather than promised in prose.

Timestamptsstring2026-08-04T14:03:11Z
When the request happened, ISO 8601.Everything is bucketed by hour. A timestamp more than an hour in the future is rejected as clock skew.
API key fingerprintkey_fpstring9f2b41c7e0a85d6314bb07e29c4af1d8
HMAC-SHA256 of the API key under your own salt, truncated to 128 bits and hex-encoded.This is the field that makes the product possible without the key itself. The server rejects anything that is not 32 hex characters, so a real key cannot be sent by accident.
Endpointendpointstring/v1/chat/completions
Which route was called.A route template, not a live path. Ids in the path are your data, not ours — the validator rejects the ones it can recognise, and your terms require templates.
Input tokenstokens_ininteger1180
Prompt tokens.A count, never the content. There is no field on this list that can carry a prompt.
Output tokenstokens_outinteger260
Completion tokens.Same as above: the size of the response, never the response.
Latencylatency_msinteger940
How long the request took.Part of the per-key usage shape. Not currently used by the fan-out rule.
HTTP statusstatusinteger200
The HTTP status you returned.Distinguishes a key being used from a key being rejected.
Client address fingerprintip_fpstring3c81a7be55d0f42a9e6c1b7380dd24f5
Fingerprint of the caller's IP address, same construction as key_fp.Counting distinct IPs per key per hour is the core of the rule. The address itself never leaves you.
Network fingerprintip_net_fpstringb04e9d217fc6a3358821e05c7ab4dd9f
Fingerprint of the surrounding network — /24 for IPv4, /64 for IPv6.Distinguishes forty addresses inside one allocation from forty scattered across the internet. The v6 prefix differs because a single /64 is one ordinary customer assignment holding 2^64 addresses.
Network block fingerprintip_block_fpstring7e15c8a0439bd6f2a1c93e470b28df61
Fingerprint of the wider block — /16 for IPv4, /48 for IPv6.The second, coarser dispersion measure. Together with ip_net_fp it is what separates a real fan-out from one noisy datacentre.
Address familyip_familystringv4
Either "v4" or "v6".The prefix sizes above are not comparable between families, so the family has to travel with them. Everything downstream is computed per family.
Costcostnumberoptional0.0041
What the request cost you, if you want to send it.The only optional field on the list, and the only one Cerberus does not need.

There is no field for a prompt, a response, an API key, or an IP address. The four fingerprints are HMAC-SHA256 under a salt that stays on your side and is never sent, so Cerberus cannot reverse them — not for you, and not for anyone who compels us.

You’ll be watching your API in about five minutes

Four steps, no call, nobody in the loop. Cerberus sits beside your API, not in front of it — nothing here is on your request path, and a Cerberus outage cannot slow or break your own service.

  1. Install it

    pip install cerberus-keys
  2. Add it to your app

    from cerberus_keys import Cerberus
    cerberus = Cerberus(ingest_token=TOKEN, salt=SALT, endpoint_url=URL)
    cerberus.record(api_key=key, ip=request.client.host,
    endpoint='/v1/chat/completions',
    tokens_in=1180, tokens_out=260, latency_ms=940, status=200)
  3. Send one request, and check it landed

    Run your app and make a single API call. Then ask Cerberus what it understood:

    curl -s -H "Authorization: Bearer $CERBERUS_TOKEN" \
    https://api.cerberushq.dev/v1/verify

    events_received counts what arrived, and verdict is a plain sentence about what we read from it. This is the step that tells you it is real — a 202 only proves bytes arrived.

  4. Connect Slack, so alerts can reach you

    Until this is set, Cerberus watches your keys and has nowhere to tell you — which looks exactly like a quiet week.

    curl -X POST https://api.cerberushq.dev/v1/webhook \
    -H "Authorization: Bearer $CERBERUS_TOKEN" \
    -H 'Content-Type: application/json' \
    -d '{"slack_webhook_url":"https://hooks.slack.com/services/..."}'

That’s it — Cerberus is watching your API. Nothing else to do: we’ll message you in Slack if a key starts looking stolen.

Detection needs about a week of each key’s own history before it can tell unusual from normal, so the first days are quiet by design. Ask us anything.

How it works, and the LiteLLM path

Why you pass the real key and the real address

Both are fingerprinted on your machine, under a salt Cerberus never receives, before anything leaves the process. That is why the sample passes ip= rather than a hash: the SDK does the work, so we cannot reverse what we are sent even in principle. Keep the salt wherever your other secrets live — it is shown once, and losing it resets baselines rather than breaking anything.

On the LiteLLM proxy instead

litellm_settings:
callbacks: cerberus_keys.litellm.handler
general_settings:
use_x_forwarded_for: true # required, not optional — see below

Both blocks are required, and the second one is the one people miss. Without use_x_forwarded_for, LiteLLM reports the socket peer — so behind a load balancer every request arrives with the balancer’s address, the distinct-IP count is permanently one, and the detector is inert while looking perfectly healthy. Cerberus checks for this shape and tells you, but it is far better not to be in it.

What /v1/verify is for

It is not a green checkmark. A wrong salt, a timestamp in local time, or a live path sent where a route template belongs all produce a cheerful 202 and then attribute nothing correctly — and you would find out a week later when no baseline exists. Verify echoes back what was actually understood, so a mistake is visible in the first minute instead of the seventh day. Call it as often as you like.

How it works

Three stages, and the third one only happens when the second is unambiguous.

  1. 1. Baseline

    40 active hours

    For each key, Cerberus learns the ordinary shape of a week: how many distinct IPs use it in an hour, and how much each of them does. It will not judge a key until that key has at least 500 requests across 40 active hours from 3 or more addresses, inside a rolling 7 days. Nothing fires while it is still learning.

  2. 2. Deviation

    4 IPs → 40 IPs

    A key starts appearing from many more places than usual, and each of those places is doing far less than usual. Five conditions must hold at once, across several consecutive hours.

  3. 3. One alert

    1 page, not 40

    One message, with the fingerprint, what changed, what normal was, and how long it has held. You acknowledge it and that key stays quiet. There is no dashboard to check.

What the rule sees

One key, one week. There is no dashboard in Cerberus — this is the shape that produces an alert, not a screen you log into.

Distinct IPs per hour

41

in the last hour, against a baseline of 4. Sustained for three consecutive hours, with requests-per-IP falling as the count rose.

Peak
41
Low
3
Average
6.3

Forty addresses each making one request is not forty times the traffic — it is roughly the same traffic, spread out. That is the signature, and it is why a busy key is not automatically a suspicious one.

What you would do instead

Most teams already have something in place. Usually it answers a question next to this one rather than this one. Cerberus is a layer, not a replacement — keep all of these.

Secret scanning on your repositories

Did this key ever appear in a commit?

A good question, and a different one. Scanning sees code you host. It does not see a key resold privately, phished from an employee, lifted off a compromised CI runner, pasted into a support ticket, or copied out of a customer's own logs. In every one of those the key is working perfectly and has never touched a public commit. Scanning tells you a key was exposed; it cannot tell you a key is being used by someone else right now.

Rate limits per key

Abuse gets capped before it costs much.

Rate limits bound the damage and never report the cause. They are a ceiling, not a detector: nothing about hitting one says the traffic is not yours, and nothing about staying under one says it is. A stolen key used carefully never trips a limit at all — five requests an hour from a network you have never seen is invisible to a limit built for a hundred. You still need the limits. They answer “how much”, and this answers “who”.

Your own logs

It is all in there already.

It is, which is exactly the problem: the signal exists and nobody is looking at it per key, per hour, against that key’s own history. Building the comparison is a week; keeping it calibrated so it does not page you every deploy is the part that does not end.

Noticing on the bill

A spike is obvious.

It is, once a month, after it has been paid. And a leaked key shared across many hands often is not a spike — the same total work, spread wider, which is the shape this looks for and an invoice cannot show.

None of the above is wrong, and none of them is watching for one key being used from many places at once. That is the only thing Cerberus claims to see, and its known limits are on this page too.

Pricing

Priced by monitored keys, per month — not by request volume, because what you are asking us to watch is keys. Nothing on this page takes payment. Cross the free limit and we get in touch to arrange billing — detection and alerting never stop while that happens.

Monitored keysPriceWho it is for
1 – 25FreeFreeFor a team watching its own keys.
26 – 100Starter$49per monthA company whose whole fleet is its own.
101 – 500Team$149per monthA small platform, or a large engineering org.
501 – 2,500Business$399per monthA platform monitoring its customers' keys.
2,501 – 10,000Scale$899per monthKeys belonging to thousands of end customers.
10,001 – 50,000Platform$1,900per monthKey monitoring as part of what you sell.
50,001 and upAbove thatTalk to usHundreds of thousands of keys, or millions.

Going over the free limit never turns detection off.

Cross the free limit and Cerberus keeps monitoring every key and keeps sending abuse alerts. Nothing about detection changes; we get in touch and arrange billing from there. Detection is a safety function, and degrading it because of a commercial threshold would mean the cheapest way to become unprotected is to grow. That is ruled out in code, not promised here.

What every band includes

  • Every detection v1 has
  • Slack alerts
  • Backfill from your own history
  • Email support — Starter and up
  • Priority response — Team and up
  • Named contact — Business and up
  • Ingest volume and retention agreed with you — Above that and up

Get an invite code

Instantly, with nobody in the loop. You redeem it in one request: the ingest token and your salt come back immediately, and the salt is shown once and never stored by us.

We use the address to know who is asking and to reach you if your integration looks wrong. Rather talk to a person first? hello@cerberushq.dev.

Known limits

Cerberus is deliberately conservative, which means there are things it does not catch, and one thing it deliberately is not. These are the ones worth knowing before you install it.