ratelimit.limit or attach it to a key. Here’s how requests are counted, what the response means, and how exact the counts are.
How requests are counted
A limit islimit requests per duration milliseconds. Unkey uses a sliding window, so a caller can’t use a full limit at the end of one window and another full limit right after. It counts the current window plus part of the previous one, based on how much of the current window is left.
Make some requests cost more
Each check spendscost tokens, 1 by default, so a heavy operation can use more of the limit than a light one. With limit: 100, a caller can make 100 checks at cost 1, 20 at cost 5, or any mix. A cost of 0 reads the current state without spending anything.
When one call checks several limits, a request denied by one limit doesn’t use up the others.
How exact the counts are across regions
Each region decides quickly using the traffic it has seen, and regions share counts within moments. In practice:
Rate limits protect you from bursts. They aren’t an exact count. Use credits when you need an exact total. See Credits and refill.
What you get back
Trust
success as the answer. Treat remaining as a close estimate, because another region may have accepted requests that aren’t counted yet.
What happens on errors
ratelimit.limit and ratelimit.multiLimit reject an empty identifier, a limit below 1, or a duration under 1000 milliseconds. If the check can’t run, they return an error, and your code decides whether to allow or deny.
keys.verifyKey is different. If the rate limit check can’t run, the verification carries on as if the limits passed, so an outage can’t lock your users out. The same happens with an inline limit below 1 or a duration under 1000 milliseconds, so check those values before you send them. See Verifying keys.