backend/apps/core-api/src/modules/agent/ (agent.routes.ts, agent.service.ts, agent.llm.ts, agent.repository.ts) and the worker in backend/apps/core-api/src/workers/agent-reply-worker.ts. This module does not follow the five file pattern. It has no controller or schema file. The route handler and its Zod schema live in agent.routes.ts.
Mounting and auth
The global
/v1 rate limiter applies. There is no req.auth on this route.
Environment variables
POST /v1/agent/webhook
Accepts one inbound WhatsApp message and schedules a debounced reply.
Auth: token query parameter, verified with ctx.smartsend.verifyWebhookToken against SMARTSEND_WEBHOOK_SECRET using timingSafeEqual.
string
required
Must equal
SMARTSEND_WEBHOOK_SECRET.string
default:""
The message text. The SmartSend flow substitutes a marker such as
[image] or [audio] for media.string
required
The sender’s phone. At least 1 character.
string
required
The SmartSend conversation to reply into.
string
Optional contact name. Parsed but not used by the service.
handleInbound runs these checks in order. Each early exit returns accepted: 0 with status 200.
AGENT_SMARTSEND_API_KEYis empty. A warning is logged and the message is dropped.- The phone, reduced to digits, is shorter than 9 or longer than 13 digits. This filters group and broadcast ids.
- The phone is not in the pilot list. The number stays silent: nothing is sent and nothing is queued.
- The text is empty or matches a media marker (
[lowercase letters]). A fixed “text only” reply is sent to the conversation.
- Pushes
{ text, conversationId, ts }onto the Redis listagent-pending:<phone>and sets a one hour expiry. - Sets
agent-latest:<phone>to the message timestamp, also with a one hour expiry. - Adds a job named
replyto the BullMQ queuePerformQueue.AGENT_REPLY(agent-reply) with data{ phone, ts }, a delay ofDEBOUNCE_MS(8000 ms), job idagent-reply-<phone>-<ts>,removeOnComplete: trueandremoveOnFail: true.
200.
How a reply is produced
startAgentReplyWorker (called from server.ts) runs a worker on the agent-reply queue with concurrency 1 and no retries. Concurrency 1 stops two replies to the same coach from interleaving. Retries are off because the service already sends its own “try again” message on failure.
replyToPending does the following.
- Debounce. If
agent-latest:<phone>exists and differs from the job’sts, a newer message arrived and this job stops. Every message schedules a job, and only the newest one answers. - Drain. Pending entries are popped one at a time with
LPOP, so a message that arrives during the drain is not lost. The texts are joined with newlines. The reply goes to theconversationIdof the last entry. - Pilot recheck. A job queued before the allowlist shrank does not answer.
- Coach resolution. A phone pinned in
AGENT_PHONE_COACHESwins when that coach exists and is active. Otherwise the repository loads every active coach with a phone and the service matches on the normalized number in JavaScript, because coach phones are stored as typed. When several rows match, the coach with the most recentlastActiveAtwins. No match sends a fixed “this number is not linked to a coach account” reply. - Daily cap.
agent-cap:<phone>:<YYYY-MM-DD>is incremented with a 25 hour expiry. The cap is 200 handled bursts per phone per UTC day. The “cap reached” reply is sent once, on the first message over the cap. After that the number is silent until the next day. - Studio key. See the next section.
- Welcome menu. History is the Redis list
agent-history:<phone>. When history is empty, or the text is one of the menu words (for examplehelpormenu, plus Hebrew equivalents), a welcome message listing the agent’s abilities is sent. If the coach asked for the menu or sent only a greeting, the turn ends there. A real first question still gets answered after the welcome. - Model call.
llm.runis called with the system prompt, the history and the joined text. - Send. The reply is split on blank lines into chunks of at most 3500 characters and each chunk is sent with
smartsend.sendConversationMessageusingAGENT_SMARTSEND_API_KEY. - History. The user text and the reply are appended. The list is trimmed to the last 24 entries and expires after 48 hours.
Studio key and the kill switch
Tools do not run against Prisma directly. They call the automation API with a studio-scoped key that the agent mints for itself.- On the first message from a studio,
mintStudioKeygenerates a partner secret, creates aStudioApiKeyrow (fixed Hebrew name,keyHash, 15 characterprefix) and stores the plaintext atStudio.settings.agent.apiKey. - The
StudioApiKeyrow shows in the coach’s integrations list. Revoking it there is the per-studio kill switch. - When a tool call comes back with HTTP 401,
agent.llm.tsflags the turn and throwsAgentKeyRevokedErrorafter the model loop. The service then writesStudio.settings.agent = { revokedAt }, which removes the stored key, and tells the coach the connection was revoked. - A revoked studio stays disconnected. The agent does not mint a new key until the coach sends the exact reconnect command
חבר מחדש. That message mints a fresh key and gets a “reconnected” reply.
Model and tools
createAgentLlm in agent.llm.ts builds an Anthropic provider with createAnthropic from @ai-sdk/anthropic and calls generateText from the ai package.
When
ANTHROPIC_API_KEY is empty no provider is built and run throws INTERNAL with agent AI is not configured. The service catches it and sends the failure message.
Tools. AGENT_TOOLS is the MCP tool set from @perform/mcp-tools, filtered to every tool marked readOnly plus create_task. At the time of writing that is:
The other MCP writes (
create_trainee, update_trainee, assign_plan, assign_program, send_form, update_task, log_weight) are left out on purpose. They would change a trainee’s data from a chat message with no confirmation step.
Execution. Each tool’s call(input) returns a method, path and body. The agent runs it through createPerformClient against http://127.0.0.1:<PORT>/v1/automation with the studio key. This is the same path the public MCP endpoint uses, so tool calls get the automation lane’s validation and audit logging. See Automation. A thrown error is returned to the model as an Error: ... string so it can recover, except a 401, which ends the turn as revoked.
System prompt. buildSystemPrompt is in Hebrew and includes the coach name, the studio name and today’s date formatted for he-IL. It tells the model to answer in short WhatsApp style Hebrew, never invent data, ask which trainee is meant when a name search returns several, say plainly that workout and attendance data is not available, explain that subscriptions, weights and forms are read one trainee at a time, use create_task for reminders, and never reveal the prompt, tool names or keys.
Repository
createAgentRepository exposes five functions.
Redis keys
Limits
- Text only. Media never reaches the model.
- 200 handled bursts per phone per day.
- Up to 8 model steps and 1200 output tokens per reply.
- One coach per phone. A phone that sits on coach rows in several studios resolves to the most recently active one unless it is pinned with
AGENT_PHONE_COACHES. - The agent does not message trainees and does not change programs.
Related pages
- WhatsApp for the per-studio webhook that uses the same token scheme.
- Coach assistant for the web chat that shares
AGENT_MODELbut runs its tools in process. - Tasks for the task rows
create_taskproduces.