We shipped an MCP server, @handset/mcp, so any MCP client (Claude Code, Claude Desktop, Cursor) can drive a real business phone system: send texts, place calls, read live transcripts and AI summaries, buy numbers, provision tenants. One config line and an agent has fifteen tools.
The most important code in the whole package is the part that refuses to run.
Giving an agent a phone is different
Most tools you hand an agent are reversible. It reads a file, it queries a table, it edits some code you review before it ships. If the agent hallucinates, you catch it at the diff.
A phone is not like that. send_message texts a real person. make_call rings a real phone. buy_number spends money every month. There is no diff to review after the fact, because the action already left the building the moment the tool returned. An agent with a live key and one bad inference is a text to the wrong customer, or a robocall you are now liable for under TCPA. The blast radius is the real world.
So the design question is not "how do we let the agent do everything." It is "what is the safe default when the agent is holding something that can act on strangers at your expense."
The default is a fake carrier
Handset has two kinds of API key. A live key (hs_live_…) sends real texts and places real calls. A test key (hs_test_…) runs against a simulated carrier: free, instant, no real phones, the whole message and call lifecycle simulated end to end.
For an AI agent, the test key is not a lesser mode. It is the expected mode. An agent should be able to explore every tool, build a whole workflow, and demo the result without ever touching a real person. So the MCP server treats a test key as the normal case and a live key as the exception that has to be argued for.
The refusal
Here is the code that runs before the server will start:
if (!KEY) {
console.error(
"HANDSET_API_KEY is not set. Use a test-mode key (hs_test_…) — get one at https://handset.dev/early-access",
);
process.exit(1);
}
if (KEY.startsWith("hs_live_") && process.env.HANDSET_ALLOW_LIVE !== "1") {
console.error(
"Refusing to start with a LIVE key: this server would let the agent text and call real people at real cost.\n" +
"Use your hs_test_… key (simulated carrier, free), or set HANDSET_ALLOW_LIVE=1 if you really mean it.",
);
process.exit(1);
}If you hand the server a live key, it does not warn you and continue. It prints why it is unhappy and exits. To run live you have to set a second, separate environment variable, HANDSET_ALLOW_LIVE=1, whose only job is to be a deliberate act. You cannot get to live by accident, by copy-pasting a key, or by an agent deciding it would be more helpful with real access. Live mode is a sentence you have to finish typing yourself.
Two small choices in there matter more than they look.
It is a hard stop, not a warning. Warnings are read by nobody, and that goes double for agents, which will happily narrate "note: this is live mode" and then send the text anyway. A process that has already called process.exit(1) cannot send a text. The only guardrail an agent cannot talk its way past is one that runs before the agent gets a turn.
The override is a different variable, not a flag on the key. The permission to go live is decoupled from the key itself. You can paste a live key into a config and the server still refuses, because holding the key and intending to use it live are two different facts, and the dangerous one has to be stated on its own.
The tools tell the agent the truth
The startup gate handles the catastrophic case. The everyday safety lives in the tool descriptions, because for an agent the description is the interface. So the outward-facing tools say exactly what they do, in the words the model reads before it decides:
place a click-to-call — dials connect_to (the agent's phone) first, then
to (the customer), and bridges them; both see the tenant number. Live mode
rings real phones and bills per minute — confirm with the user first. Test
mode simulates the whole lifecycle.
buy_number says it is "BILLABLE in live mode (monthly rental)… Confirm with the user before buying on a live key." send_message says it "texts a real person and bills per segment — confirm with the user first." The pattern is deliberate: every tool that costs money or reaches a stranger is labeled OUTWARD-FACING or BILLABLE, states that live mode is real, and tells the agent to check with the human first. In test mode those same tools are free and simulated, so the agent can rehearse the entire flow with the guardrails narrated but nothing at stake.
When it does start, the server says which world it is in, on stderr, every time:
handset-mcp: test mode against https://api.handset.dev/v1
No ambiguity about whether the phone in the agent's hand is loaded.
What we learned
The safest default for a tool that can act in the world is one that boots up unable to. Make the harmless mode the path of least resistance (here, a test key against a simulated carrier that costs nothing and cannot reach anyone), and make the dangerous mode a deliberate, typed, separate override that no amount of agent helpfulness can reach on its own.
It is not a clever mechanism. It is one if statement and an environment variable. But it encodes the right belief about what an autonomous agent holding a phone should be able to do by default, which is: explore everything, break nothing, and call a real person only when a human said so out loud.
If you want to give your own agent a phone system to play with, it is one line: claude mcp add handset -e HANDSET_API_KEY=hs_test_… -- npx -y @handset/mcp. It starts in test mode, against a carrier that is not real, on purpose.
More from Handset