What Cerberus receives
Twelve fields, nothing else. The ingest endpoint rejects any unknown field by name, so this list is enforced in code rather than promised in prose.
- Timestamptsstring2026-08-04T14:03:11Z
- When the request happened, ISO 8601.Everything is bucketed by hour. A timestamp more than an hour in the future is rejected as clock skew.
- API key fingerprintkey_fpstring9f2b41c7e0a85d6314bb07e29c4af1d8
- HMAC-SHA256 of the API key under your own salt, truncated to 128 bits and hex-encoded.This is the field that makes the product possible without the key itself. The server rejects anything that is not 32 hex characters, so a real key cannot be sent by accident.
- Endpointendpointstring/v1/chat/completions
- Which route was called.A route template, not a live path. Ids in the path are your data, not ours — the validator rejects the ones it can recognise, and your terms require templates.
- Input tokenstokens_ininteger1180
- Prompt tokens.A count, never the content. There is no field on this list that can carry a prompt.
- Output tokenstokens_outinteger260
- Completion tokens.Same as above: the size of the response, never the response.
- Latencylatency_msinteger940
- How long the request took.Part of the per-key usage shape. Not currently used by the fan-out rule.
- HTTP statusstatusinteger200
- The HTTP status you returned.Distinguishes a key being used from a key being rejected.
- Client address fingerprintip_fpstring3c81a7be55d0f42a9e6c1b7380dd24f5
- Fingerprint of the caller's IP address, same construction as key_fp.Counting distinct IPs per key per hour is the core of the rule. The address itself never leaves you.
- Network fingerprintip_net_fpstringb04e9d217fc6a3358821e05c7ab4dd9f
- Fingerprint of the surrounding network — /24 for IPv4, /64 for IPv6.Distinguishes forty addresses inside one allocation from forty scattered across the internet. The v6 prefix differs because a single /64 is one ordinary customer assignment holding 2^64 addresses.
- Network block fingerprintip_block_fpstring7e15c8a0439bd6f2a1c93e470b28df61
- Fingerprint of the wider block — /16 for IPv4, /48 for IPv6.The second, coarser dispersion measure. Together with ip_net_fp it is what separates a real fan-out from one noisy datacentre.
- Address familyip_familystringv4
- Either "v4" or "v6".The prefix sizes above are not comparable between families, so the family has to travel with them. Everything downstream is computed per family.
- Costcostnumberoptional0.0041
- What the request cost you, if you want to send it.The only optional field on the list, and the only one Cerberus does not need.
There is no field for a prompt, a response, an API key, or an IP address. The four fingerprints are HMAC-SHA256 under a salt that stays on your side and is never sent, so Cerberus cannot reverse them — not for you, and not for anyone who compels us.
You’ll be watching your API in about five minutes
Four steps, no call, nobody in the loop. Cerberus sits beside your API, not in front of it — nothing here is on your request path, and a Cerberus outage cannot slow or break your own service.
Install it
pip install cerberus-keysAdd it to your app
from cerberus_keys import Cerberuscerberus = Cerberus(ingest_token=TOKEN, salt=SALT, endpoint_url=URL)cerberus.record(api_key=key, ip=request.client.host,endpoint='/v1/chat/completions',tokens_in=1180, tokens_out=260, latency_ms=940, status=200)Send one request, and check it landed
Run your app and make a single API call. Then ask Cerberus what it understood:
curl -s -H "Authorization: Bearer $CERBERUS_TOKEN" \https://api.cerberushq.dev/v1/verifyevents_receivedcounts what arrived, andverdictis a plain sentence about what we read from it. This is the step that tells you it is real — a 202 only proves bytes arrived.Connect Slack, so alerts can reach you
Until this is set, Cerberus watches your keys and has nowhere to tell you — which looks exactly like a quiet week.
curl -X POST https://api.cerberushq.dev/v1/webhook \-H "Authorization: Bearer $CERBERUS_TOKEN" \-H 'Content-Type: application/json' \-d '{"slack_webhook_url":"https://hooks.slack.com/services/..."}'
That’s it — Cerberus is watching your API. Nothing else to do: we’ll message you in Slack if a key starts looking stolen.
Detection needs about a week of each key’s own history before it can tell unusual from normal, so the first days are quiet by design. Ask us anything.
How it works, and the LiteLLM path
Why you pass the real key and the real address
Both are fingerprinted on your machine, under a salt Cerberus never receives, before anything leaves the process. That is why the sample passes ip= rather than a hash: the SDK does the work, so we cannot reverse what we are sent even in principle. Keep the salt wherever your other secrets live — it is shown once, and losing it resets baselines rather than breaking anything.
On the LiteLLM proxy instead
Both blocks are required, and the second one is the one people miss. Without use_x_forwarded_for, LiteLLM reports the socket peer — so behind a load balancer every request arrives with the balancer’s address, the distinct-IP count is permanently one, and the detector is inert while looking perfectly healthy. Cerberus checks for this shape and tells you, but it is far better not to be in it.
What /v1/verify is for
It is not a green checkmark. A wrong salt, a timestamp in local time, or a live path sent where a route template belongs all produce a cheerful 202 and then attribute nothing correctly — and you would find out a week later when no baseline exists. Verify echoes back what was actually understood, so a mistake is visible in the first minute instead of the seventh day. Call it as often as you like.
How it works
Three stages, and the third one only happens when the second is unambiguous.
1. Baseline
40 active hoursFor each key, Cerberus learns the ordinary shape of a week: how many distinct IPs use it in an hour, and how much each of them does. It will not judge a key until that key has at least 500 requests across 40 active hours from 3 or more addresses, inside a rolling 7 days. Nothing fires while it is still learning.
2. Deviation
4 IPs → 40 IPsA key starts appearing from many more places than usual, and each of those places is doing far less than usual. Five conditions must hold at once, across several consecutive hours.
3. One alert
1 page, not 40One message, with the fingerprint, what changed, what normal was, and how long it has held. You acknowledge it and that key stays quiet. There is no dashboard to check.
What the rule sees
One key, one week. There is no dashboard in Cerberus — this is the shape that produces an alert, not a screen you log into.
Distinct IPs per hour
41
in the last hour, against a baseline of 4. Sustained for three consecutive hours, with requests-per-IP falling as the count rose.
- Peak
- 41
- Low
- 3
- Average
- 6.3
Forty addresses each making one request is not forty times the traffic — it is roughly the same traffic, spread out. That is the signature, and it is why a busy key is not automatically a suspicious one.
What you would do instead
Most teams already have something in place. Usually it answers a question next to this one rather than this one. Cerberus is a layer, not a replacement — keep all of these.
Secret scanning on your repositories
Did this key ever appear in a commit?
A good question, and a different one. Scanning sees code you host. It does not see a key resold privately, phished from an employee, lifted off a compromised CI runner, pasted into a support ticket, or copied out of a customer's own logs. In every one of those the key is working perfectly and has never touched a public commit. Scanning tells you a key was exposed; it cannot tell you a key is being used by someone else right now.
Rate limits per key
Abuse gets capped before it costs much.
Rate limits bound the damage and never report the cause. They are a ceiling, not a detector: nothing about hitting one says the traffic is not yours, and nothing about staying under one says it is. A stolen key used carefully never trips a limit at all — five requests an hour from a network you have never seen is invisible to a limit built for a hundred. You still need the limits. They answer “how much”, and this answers “who”.
Your own logs
It is all in there already.
It is, which is exactly the problem: the signal exists and nobody is looking at it per key, per hour, against that key’s own history. Building the comparison is a week; keeping it calibrated so it does not page you every deploy is the part that does not end.
Noticing on the bill
A spike is obvious.
It is, once a month, after it has been paid. And a leaked key shared across many hands often is not a spike — the same total work, spread wider, which is the shape this looks for and an invoice cannot show.
None of the above is wrong, and none of them is watching for one key being used from many places at once. That is the only thing Cerberus claims to see, and its known limits are on this page too.
Pricing
Priced by monitored keys, per month — not by request volume, because what you are asking us to watch is keys. Nothing on this page takes payment. Cross the free limit and we get in touch to arrange billing — detection and alerting never stop while that happens.
| Monitored keys | Price | Who it is for |
|---|---|---|
| 1 – 25Free | Free | For a team watching its own keys. |
| 26 – 100Starter | $49per month | A company whose whole fleet is its own. |
| 101 – 500Team | $149per month | A small platform, or a large engineering org. |
| 501 – 2,500Business | $399per month | A platform monitoring its customers' keys. |
| 2,501 – 10,000Scale | $899per month | Keys belonging to thousands of end customers. |
| 10,001 – 50,000Platform | $1,900per month | Key monitoring as part of what you sell. |
| 50,001 and upAbove that | Talk to us | Hundreds of thousands of keys, or millions. |
Going over the free limit never turns detection off.
Cross the free limit and Cerberus keeps monitoring every key and keeps sending abuse alerts. Nothing about detection changes; we get in touch and arrange billing from there. Detection is a safety function, and degrading it because of a commercial threshold would mean the cheapest way to become unprotected is to grow. That is ruled out in code, not promised here.
What every band includes
- Every detection v1 has
- Slack alerts
- Backfill from your own history
- Email support — Starter and up
- Priority response — Team and up
- Named contact — Business and up
- Ingest volume and retention agreed with you — Above that and up
Get an invite code
Instantly, with nobody in the loop. You redeem it in one request: the ingest token and your salt come back immediately, and the salt is shown once and never stored by us.
We use the address to know who is asking and to reach you if your integration looks wrong. Rather talk to a person first? hello@cerberushq.dev.
Known limits
Cerberus is deliberately conservative, which means there are things it does not catch, and one thing it deliberately is not. These are the ones worth knowing before you install it.