Every usage-based API is a machine for turning actions into money. Send a text, spend a fraction of a cent. Buy a number, spend two dollars a month. The machine has exactly one bug it is not allowed to have: charging you for something twice. A dashboard that loads a little slow is a Tuesday. A customer who spots a duplicate charge never fully trusts the number on the invoice again, and they are right not to.
The uncomfortable part is that double-billing is the default outcome, not an edge case. Webhooks get redelivered. Jobs get retried. A worker crashes after it charged you but before it recorded that it charged you, comes back up, and charges you again. Money correctness is not something you add; it is something you refuse to lose at every step. Here is how the prepaid wallet on Handset holds onto it.
The ledger is append-only
The wallet is not a balance column. A balance column is a number you UPDATE, and every UPDATE is a race and a place to lose an audit trail. It is one signed, append-only table:
CREATE TABLE account_credits (
id text PRIMARY KEY, -- cr_…
account_id text NOT NULL REFERENCES accounts (id),
amount_cents bigint NOT NULL, -- + credit, - debit
kind text NOT NULL
CHECK (kind IN ('topup','grant','usage','adjustment','refund')),
reference text, -- a Stripe session id, a usage rollup key, …
description text,
created_at timestamptz NOT NULL DEFAULT now()
);Positive rows are money in: a top-up, a grant. Negative rows are money out: a usage debit, an adjustment. The balance is not stored anywhere; it is a question you ask:
func (s *Service) Balance(ctx context.Context, accountID string) (int64, error) {
var cents int64
err := s.pool.QueryRow(ctx,
`SELECT COALESCE(SUM(amount_cents), 0) FROM account_credits
WHERE account_id = $1`, accountID).Scan(¢s)
return cents, err
}We never mutate a row. A refund is a new positive row, not a deleted debit. A correction is an adjustment row, not an edit. This costs almost nothing (the sum is indexed and the table is small per account) and it buys the one thing you want when a customer emails you about their balance: every cent has a row, a timestamp, and a reason.
Idempotency is a unique index, not application logic
The tempting way to prevent double-charges is to check first: has this already been billed? no? then bill it. That check-then-act is two statements with a gap in the middle, and two workers will both read "no" in that gap and both bill it. The database is the only thing that can make "charge at most once" atomic, so we let it:
CREATE UNIQUE INDEX account_credits_reference_uniq
ON account_credits (account_id, kind, reference) WHERE reference IS NOT NULL;Every charge that must happen at most once carries a reference that identifies the thing being charged for, and the insert either succeeds or hits the unique index. There is no read:
func CreditTx(ctx context.Context, tx pgx.Tx, accountID string, amountCents int64,
kind, reference, description string) (string, error) {
id := ids.New(ids.Credit)
_, err := tx.Exec(ctx, `
INSERT INTO account_credits (id, account_id, amount_cents, kind, reference, description)
VALUES ($1, $2, $3, $4, $5, $6)`,
id, accountID, amountCents, kind, nullIfEmpty(reference), nullIfEmpty(description))
if err != nil {
if isUniqueViolation(err) {
return "", ErrDuplicateReference // already charged; not an error to the caller
}
return "", fmt.Errorf("billing: credit: %w", err)
}
return id, nil
}ErrDuplicateReference is the whole design. It is not a failure; it is the system telling you this exact charge already landed, and the caller's correct response is almost always to shrug and move on. A Stripe top-up uses the checkout session id as its reference, so a webhook Stripe redelivers three times credits the account once. The other two inserts bounce off the index and return ErrDuplicateReference, and the handler returns 200 without touching the balance.
The index is scoped to (account_id, kind, reference), not to reference alone. A reference is only meaningful inside one account's ledger, and one account's checkout id should never be able to collide with another's. We learned that the boring way: an earlier version keyed on (kind, reference) globally, and a test account's rollup reference could, in principle, block a real account's. Scoping it to the account is both more correct and more obviously correct.
Usage debits: price the rollup, stamp it once
Top-ups are easy because Stripe hands you a natural idempotency key. Usage is harder, because usage is a firehose of individual events (each sent segment, each voice minute) and you cannot afford a ledger row per event. So usage is aggregated into hourly rollups first, and a sweep prices the finalized ones:
func (s *Service) SweepWalletDebits(ctx context.Context, limit int) ([]string, error) {
cutoff := time.Now().UTC().Truncate(time.Hour).Add(-FinalizeAfter)
rows, _ := s.pool.Query(ctx, `
SELECT r.account_id, a.pricing_tier, r.kind, r.hour, r.quantity
FROM usage_rollups r
JOIN accounts a ON a.id = r.account_id
WHERE r.wallet_debited_at IS NULL
AND a.billing_mode = 'prepaid'
AND r.hour < $1
ORDER BY r.hour LIMIT $2`, cutoff, limit)
// …Two things make this safe to run forever, on a timer, with retries, and never double-charge.
First, it only touches rollups older than FinalizeAfter (a few hours). A rollup for the current hour is still accumulating; pricing it now would bill half a message. The hour < cutoff clause means we only ever charge for time that is closed.
Second, the debit and the "I charged this" mark happen in one transaction, and there are two independent guards against a repeat:
ref := fmt.Sprintf("%s:%s:%d", p.accountID, p.kind, p.hour.Unix())
cents := int64(math.Round(p.quantity * priceFor(p.tier)[p.kind]))
err := s.store.InAccountTx(ctx, p.accountID, func(tx pgx.Tx) error {
if cents > 0 {
_, e := billing.CreditTx(ctx, tx, p.accountID, -cents, "usage", ref, desc)
if e != nil && !errors.Is(e, billing.ErrDuplicateReference) {
return e // a real error; a duplicate is fine
}
}
_, e := tx.Exec(ctx, `
UPDATE usage_rollups SET wallet_debited_at = now()
WHERE account_id = $1 AND kind = $2 AND hour = $3`,
p.accountID, p.kind, p.hour)
return e
})The reference is account:kind:hour, so the same rollup can only ever produce one debit row through the unique index, no matter how many times the sweep runs. And the rollup gets stamped with wallet_debited_at, so the next sweep's WHERE wallet_debited_at IS NULL filters it out before it is even considered. Either guard alone would be enough. Having both means a crash between the debit and the stamp is harmless: the stamp is missing so the row is reconsidered, but the debit's reference already exists, so the re-attempt returns ErrDuplicateReference and the transaction just re-applies the stamp. Belt and suspenders, because the failure mode is someone's money.
Two billing modes that cannot both fire
Not every account is prepaid. Larger customers get invoiced monthly against Stripe meters. That is a second path from usage to money, and two paths from the same usage is exactly how you double-bill an entire customer. So the mode is a column, and the two paths are written to be mutually exclusive by construction:
ALTER TABLE accounts ADD COLUMN billing_mode text NOT NULL DEFAULT 'prepaid'
CHECK (billing_mode IN ('prepaid', 'invoiced'));The wallet sweep you saw above filters billing_mode = 'prepaid'. The Stripe meter sync does the opposite and filters billing_mode = 'invoiced'. A given account's usage is eligible for exactly one of them, ever. There is no shared code path where a flag might be read wrong; the two workers select disjoint sets of accounts out of the same table.
The migration that almost back-billed everyone
When prepaid billing shipped, the account_credits table was new but usage_rollups was not. Accounts had months of usage sitting in rollups with a null wallet_debited_at. The very first sweep would have looked at all of it, priced it, and charged every account for its entire history in one run. That is not a bug in the sweep; the sweep did exactly what it says. It is a bug in when the wallet started existing.
The fix was one line in the migration that introduced the column:
ALTER TABLE usage_rollups ADD COLUMN wallet_debited_at timestamptz;
-- Everything that predates the wallet is already settled. Without this,
-- the first sweep back-bills every account's entire history.
UPDATE usage_rollups SET wallet_debited_at = now() WHERE wallet_debited_at IS NULL;Stamp all existing rollups as already-debited, so the sweep only ever charges usage incurred after prepaid billing went live. It is a two-word idea (already settled) and it is the difference between a launch and a mass of refund emails. The lesson we took: whenever you introduce a system that acts on existing history, the migration is not done until you have decided, explicitly, what that system believes about everything that happened before it existed.
Enforcement lives before the carrier, not after
One last piece, because a wallet is only real if it can say no. A live send checks the balance before it touches a carrier:
func (s *Service) CanSpendLive(ctx context.Context, accountID string) (bool, error) {
// prepaid accounts must have a positive balance; invoiced accounts always may
// …
}If a prepaid account is at zero, the send never leaves the building. It returns 402 insufficient_balance with a code the caller can branch on, and nothing is queued to a carrier that we would then have to unwind. Test-mode sends skip this entirely, because test usage is free and the simulated carrier costs us nothing. The check is cheap (it is that same indexed SUM) and it runs on the request path, so the wallet is a gate, not a report you read after the money is gone.
What we learned
Money correctness is not a feature you build once. It is a property you decline to give up at every layer: the ledger refuses to mutate, the index refuses a second charge, the sweep refuses to price an open hour, the two billing modes refuse to overlap, the migration refuses to back-bill, and the send refuses to leave without a balance. None of those are clever. The discipline is in making each one boring and letting the database, not your application code, be the thing that enforces "at most once." Application logic has gaps between statements. A unique constraint does not.
The nicest outcome is the one nobody sees: partners send us usage all day, the sweep runs on its timer, webhooks redeliver, workers retry, and the number on the wallet is exactly the sum of what happened. No duplicate charges, no back-bills, and a row you can point at for every cent.
More from Handset