Engineering Case study

Designing an Operations Heartbeat System

A system design write-up on turning UI activity events into reliable agent state across compute windows, HCM state commits, downstream propagation, multi-region cutover, and service split boundaries.

Published 16 min read 中文版

An operations platform starts with a simple question: how do we know whether an agent is online?

The smallest version is almost boring. The frontend sends a heartbeat every few seconds. The backend stores last_seen_at. If the timestamp is too old, the agent becomes offline.

That version can light up a green dot. It is nowhere near enough for dispatching, workforce management, routing, monitoring, and time tracking. In the system I worked on, the HCM and Heartbeat path served 47,149 agents, 7,000 skill groups, about 4,000 QPS, and 141 upstream services. At that scale, “online” becomes a state platform.

Here is the target shape first.

Final target HCM heartbeat state platform
Figure 0: The final target is a state platform. User activity enters through the Workbench SDK, WS-API, and Frontier long-link path, then flows through Heartbeat, MQ, Compute, HCM, and downstream state propagation. Generated by gpt-image-2.

The rest of this article builds the system from last_seen_at to that final design, one engineering problem at a time.

Can we tell whether an agent is active?

Version 0 writes state directly.

Version 0 direct status write
Figure 1: The naive heartbeat design records the latest heartbeat timestamp. It has no task count, skill group context, business rules, or region ownership. Generated by gpt-image-2.

The data model is tiny:

agent_id
last_seen_at
status

The frontend reports heartbeat events. The backend updates last_seen_at. A scheduled job scans expired records and marks agents as OFFLINE.

This model works for a lightweight presence indicator. Before we even talk about task count or skill groups, the intake path already has problems:

  • A workbench page has multiple modules. They should not each open their own channel.
  • User activity is bursty. A typing burst should not become a synchronous write storm.
  • Browser connections drop. A reconnect should not make the system lose recent activity.
  • Long-link push and short-poll compensation can deliver the same logical message twice.
  • Heartbeat service restarts should not erase the input stream.
  • State calculation should be deployable without changing the browser SDK.

last_seen_at gives us a timestamp. The next problem is to turn browser activity into a reliable event stream.

How do UI events enter Heartbeat reliably?

The next layer moves activity intake out of direct writes and into an event stream.

Version 1 event intake
Figure 2: Workbench activity enters through SDK, WS-API, and Frontier before Heartbeat writes the raw event stream to MQ. State calculation comes later. Generated by gpt-image-2.

The frontend does more than send a periodic heartbeat. It reports clicks, key events, mouse activity, URL changes, ticket open, ticket reply, ticket completion, handoff, escalation, and other workbench actions.

These events mean different things, but they share one property: they prove an agent did something at a specific time.

Between the agent UI and Heartbeat, the workbench uses a long-link path:

Agent UI
-> Workbench SDK
-> WS-API
-> Frontier
-> Heartbeat Service
-> Raw MQ

The Workbench SDK keeps one long-link connection per window. Business modules register by module + entity, and deviceID identifies the channel. That granularity matters. If entity is too fine, a user switching entities can hit a short unregistered window and lose a push.

WS-API handles long-link initialization, module registration, deregistration, and short-link compensation. Frontier owns the long-link channel. When Frontier has problems, the client can use short polling to fetch messages that were not acked. The default interval is 30 seconds and can be lowered by configuration during an incident.

We also dedupe early. Business messages carry ReqID, and the SDK keeps a recent ReqID cache. The RPC push service generates msgID, so long-link delivery and short-poll compensation can dedupe the same logical message. If a business flow needs ordering, it carries an Index; ordering semantics stay with the business.

Each node earns its place:

Node / changeProblem it solvesSoftware engineering idea
Workbench SDKMultiple modules in one page need one shared client-side channelConnection pooling and client-side ownership
module + entity registrationEvents must be routed to the right business module without opening new linksNamespacing and subscription boundaries
deviceIDA backend push target needs a stable browser-window identitySession identity
WS-APIInit, register, deregister, and compensation need one control surfaceControl plane separate from data delivery
FrontierLong-link delivery should be handled by a dedicated channel serviceSeparation of concerns
Short-poll compensationLong-link outages should not drop recent messagesGraceful degradation
ReqID / msgID dedupePush and compensation can deliver the same logical event twiceIdempotency
Raw MQBursty activity and service restarts should not hit Compute directlyBackpressure and durable buffering
Heartbeat intake serviceBrowser protocol details should not leak into ComputeAdapter boundary

The queue contains input, not truth.

If a consumer turns every keyup into ONLINE, and every missing event into ABNORMAL, the system will still misclassify people. Events lack task count, skill group membership, work status, rule scope, and notification history.

That leads to the next set of questions, which are about business context: can this agent receive a ticket, which skill groups does the rule cover, has the abnormal status already been handled, and which region can emit a state change? Compute and HCM own those decisions.

That leads to the compute layer.

How do events become state candidates?

Compute turns raw activity into candidate state changes.

Version 2 compute layer
Figure 3: Compute consumes raw events, reads current facts from HCM, and produces candidate status changes. Generated by gpt-image-2.

Compute has one job: convert “an event happened” into “this agent may need a state change.”

Event types fall into a few groups:

TypeMeaning
1Online heartbeat
2Work status change
3-12Ticket opened, replied, completed, transferred, escalated
100-105Click, key event, mouse, status switch, outbound call, URL change
3001No action for 15 minutes result
3002Abnormal for 15 minutes result
3003Not on app
4000Rule Config V2 result

Compute does not write the final status. It first asks HCM for current facts:

GetWorkStatus(agent_id, tenant_id, channel)

HCM returns current work status, task count, skill group membership, and related facts. Compute combines those facts with the rule configuration:

event + current facts + rule = candidate

A candidate is still only a candidate. An agent may have no action after 10:00, so Compute may produce an abnormal candidate at 10:10. At 10:10:01, the agent may receive a new ticket. HCM must re-read the facts before committing the state.

Keep that boundary hard: Compute calculates candidates. HCM commits facts.

Version 2 adds these pieces for specific reasons:

Node / changeProblem it solvesSoftware engineering idea
ComputeRaw events need business interpretation before they can affect stateDomain service boundary
GetWorkStatus readA candidate needs current task, status, and skill group factsRead-before-decide
Candidate state changeCalculation should not write final truth directlyCommand staging
Rule configurationStatus logic changes faster than service codePolicy/data separation
HCM recheck requirementFacts can change after Compute produced a candidateOptimistic validation

How do we calculate “no action for N minutes”?

Many heartbeat rules are time-window rules:

ScenarioSystem action
ONLINE and no task for 30 secondsMove to IDLE
ONLINE / IDLE / BUSY and no action, or not on app, for 8 minutesSend a reminder
ONLINE / IDLE / BUSY and no action, or not on app, for 10 minutesMove to ABNORMAL
ABNORMAL and still no action, or not on app, for another 10 minutesMove to OFFLINE

The obvious implementation is one timer per agent.

That becomes painful quickly. Timers disappear on process restart. Autoscaling spreads timers across instances. Region cutover now has to move timer ownership. For tens of thousands of agents, in-process timers couple business state to instance lifetime.

We used Redis zset to represent time windows.

Version 3 Redis zset time window
Figure 4: Candidate agents are written into Redis zsets by event timestamp. Cron scans due records and sends exception messages back into the state update path. Generated by gpt-image-2.

Redis keeps several queues:

agent_no_action_for_8_min
agent_no_action_for_10_min
queue_agent_no_action_for_15_min
agent_in_abnormal_status_more_then_10_min
queue_agent_in_abnormal_status_more_then_15_min
agent_not_on_app_for_5_min
queue_agent_not_on_app

The zset score is the timestamp. A due scan uses zrangebyscore for everything older than now - threshold.

This data structure fits the problem:

Needzset behavior
Time-window checksScore is timestamp
Compute restart safetyCandidates stay in Redis
Multiple rulesSeparate zsets per rule family
Duplicate reminder controlAdd agent/message/rule locks
Canary and rollbackQueue, rule, and IDC switches stay configurable

Cron scans every 2 seconds. Before emitting messages, it checks whether the current IDC can send and grabs a short TTL global lock to avoid duplicate emission from multiple instances.

One detail matters more than it looks: the business window and the dedupe lock window should be different.

For a 10-minute no-action reminder, the business threshold is 600 seconds. The agent-level dedupe lock can be 840 seconds. That reduces reminder spam while still leaving room for later transitions such as ABNORMAL -> OFFLINE.

Version 3 adds these pieces:

Node / changeProblem it solvesSoftware engineering idea
Redis zsetTens of thousands of timers should not live inside process memoryExternalized state
Timestamp scoreDue candidates need efficient time-window scansIndex by time
Separate rule queuesDifferent windows and rules should not share hidden stateWork partitioning
Cron scannerDue work needs a repeatable executor outside request flowScheduled worker
Global TTL lockMultiple instances may scan the same queueLease-based coordination
Dedupe lock windowRepeated reminders should be bounded independently from rule thresholdIdempotency window

Who writes the final status?

Compute emits candidates. HCM writes the final status.

HCM owns the agent, skill group, task count, current status, and state-change log. If several services can write status independently, downstream systems will see competing facts.

The status enum looks like this:

ONLINE(1000)
TRAINING(1001)
BREAK(1002)
MEETING(1003)
ABNORMAL(1004)
IDLE(1005)
LUNCH(1006)
OFFLINE(1007)
OTHER(1008)
BUSY(1009)
STANDBY(1010)

When HCM consumes an exception candidate, it reads the current facts again and checks:

  1. Does the current status still match the rule?
  2. Does the current task count allow this transition?
  3. Is the agent in a covered skill group?
  4. Is this region allowed to send state changes?
  5. Has this exception already been notified or handled?

A rule can look like this:

{
"access_party_ids": [2, 3, 9, 45, 46],
"status": [1000, 1005, 1009],
"no_heart_time_limit": 600,
"status_to": 1004,
"status_change_note": "No action for 10min, automatically changes to abnormal.",
"notifies": [
{
"type": "Lark",
"title": "abnormal hint"
}
]
}

The rule affects both the state transition and the notification. Compute configuration, Redis queues, and HCM rules all constrain the final behavior.

HCM’s responsibilities stay small and strict:

  • Reject stale candidates.
  • Recheck current facts.
  • Write status_table.
  • Write unified_work_status_log.
  • Hand the state change to the downstream propagation path.

Version 4 introduces a state authority:

Node / changeProblem it solvesSoftware engineering idea
HCM as writerMultiple writers would create competing status factsSingle writer / source of truth
Status enumCallers need stable state semanticsExplicit domain model
Rule coverage checkRules apply to selected status, skill group, and business scopePolicy enforcement
Task-count recheckWork assignment can change after Compute emitted the candidateConsistency at commit time
Status logState changes need auditability and downstream replay contextAppend-only history

How do downstream systems get the same state?

After HCM commits a status, WFM, routing, analytics, and other consumers still need the same fact.

Version 4 final single-region pipeline
Figure 5: The single-region pipeline splits intake, candidate calculation, time windows, state commit, and downstream propagation. Generated by gpt-image-2.

The propagation path is:

HCM UpdateWorkStatus
-> status_table
-> DBus / binlog
-> status change MQ
-> WFM / Routing / Analytics

This gives consumers a DB-backed fact stream. The cost is a longer path. A slow binlog handler, event dispatcher, MQ backlog, or downstream consumer can make users see stale state.

We hit that failure mode.

During one US-region incident, WFM’s omni-channel view did not show the current status. HCM DB already had the agent offline, and the HCM-to-WFM send path looked successful.

The real delay was earlier. The binlog-to-HCM-MQ path had about 300k messages queued. The work status handler was not the slow part. Another handler that depended on ES was timing out, and multiple handlers shared one consumption path. The slow handler dragged status propagation with it.

The fix direction was obvious after the incident: add handler-level latency metrics, and move work status propagation closer to a direct binlog-to-RMQ path so unrelated event distribution cannot block it.

The design question is simple: which handler can stall the state path, which consumer group shares a failure domain, and which metric proves the downstream view received the state?

The propagation part of Version 4 adds a second set of boundaries:

Node / changeProblem it solvesSoftware engineering idea
status_tableDownstream systems need a committed fact, not a candidateDurable source of truth
DBus / binlogConsumers need to follow DB commits without coupling to write RPCsChange data capture
Status change MQWFM and routing should consume state asynchronouslyEvent-driven propagation
Handler metricsA slow handler can hide behind a successful DB writeObservability by stage
Consumer isolation planUnrelated handlers can block work status propagationFailure-domain isolation

How do events survive a region cutover?

The single-region pipeline works until disaster recovery and traffic cutover enter the picture.

Heartbeat events are the input to state calculation. If traffic moves to a target region before that region has recent Raw MQ events, Compute loses the time-window context. Agents can be marked no-action right after cutover, or abnormal calculation can pause until enough fresh context accumulates.

The first cross-region change is MQ mirror.

Version 5 cross-region MQ mirror
Figure 6: Raw MQ is mirrored between regions so the target region already has a consumable activity stream before it receives traffic. Generated by gpt-image-2.

Mirror protects event continuity:

Region A Raw MQ <-> Region B Raw MQ

Before cutover, both regions can see recent heartbeat events. When traffic actually moves, the target region is not starting from an empty queue. It already sees recent clicks, key events, heartbeats, ticket actions, and status events.

The cutover runbook then follows the data path:

  1. Check Heartbeat, HCM, Raw MQ, binlog MQ, and status MQ in the current region.
  2. Move a small traffic slice, for example 2%.
  3. Check from_dc dimensions to verify the target region receives expected traffic.
  4. Check mirrored Raw MQ lag, error rate, and consumption rate.
  5. If anything looks wrong, remove the routing config and cut traffic back.

MQ mirror reduces event loss. It also creates a new problem: two regions can see the same event.

Duplicate events are tolerable. Duplicate state changes are dangerous. The same no-action candidate emitted in two regions can send two exception messages to HCM, and downstream systems may receive repeated state changes.

That takes us to dedupe.

Version 5 adds event continuity across regions:

Node / changeProblem it solvesSoftware engineering idea
MQ mirrorTarget region needs recent activity before it receives trafficReplicated log
from_dc dimensionOperators need to prove where traffic and events are flowingTraceable provenance
Lag and consume-rate checksCutover should wait for the target stream to catch upReadiness gate
Small traffic sliceRegion changes need a reversible first stepCanary cutover
Rollback by routing configCutover failure should not require code rollbackOperational control plane

How do we handle duplicates after mirror?

Duplicate control has two stages.

The sender side ensures only one region emits exception messages.

The state side makes HCM re-read current facts and apply idempotent checks before writing.

Version 6 region dedupe
Figure 7: Mirror keeps events continuous, while active-region ownership and dedupe locks keep state changes single-writer. Generated by gpt-image-2.

Compute can consume mirrored events in both regions. Cron checks active_idc before emitting due candidates:

active_idc = true -> can emit exception message
active_idc = false -> calculate only, no emit

Before sending, Cron also takes a dedupe lock. The lock must match the state transition boundary. Too broad, and unrelated rules block each other. Too narrow, and repeated messages for the same stage leak through.

A practical lock key is:

agent_id + rule_id + window_start

Redis zset helps collapse part of the duplicate input, because the same member written twice updates the score. The actual send path still needs the agent/rule/window lock.

HCM remains the final guard. When it receives an exception message, it re-reads current status, task count, skill group membership, and rule config. It writes status_table and unified_work_status_log only if the current facts still match the rule. If the agent already has a task, or the status changed through another path, HCM drops the candidate.

After cross-region support, the state path has three independent switches:

SwitchJob
Traffic routingControls which region receives requests
MQ mirrorControls whether events copy across regions
active_idcControls which region can emit exception status messages

Keep those switches separate. Request routing, event replication, and state emission ownership often need different rollout and rollback timing.

Version 6 turns mirrored events into safe state changes:

Node / changeProblem it solvesSoftware engineering idea
active_idcTwo regions can calculate, but only one should emit state changesLeader ownership
Agent/rule/window lockThe same due candidate can be seen more than onceIdempotency key
HCM recheckDuplicate or stale candidates should not commit stale stateDefensive write validation
Separate traffic / mirror / emit switchesRequest flow, event replication, and write ownership change at different speedsOrthogonal control planes

How do we split TT from the main business line without breaking the old path?

The system later served both the main business line (labelled Main below) and TT. Sharing HCM, Heartbeat, DB, MQ, and routing update paths tied release risk, capacity, and disaster recovery together.

The tempting split is to create TT HCM, TT Heartbeat, TT DB, and TT MQ, then move all TT upstreams in one shot.

That is a bad bet. HCM touches login, status update, skill groups, task count, routing, WFM, and binlog propagation. A field mismatch, filter bug, or consumer group mistake can affect dispatch and state sync.

The safer first phase is coexistence.

Version 7: split TT with a coexistence phase
Figure 8: TT upstreams first move to TT HCM, while writes still land in Main HCM. Forwarding and dsyncer leave room for rollback. Generated by gpt-image-2.

The phase-one path is:

TT upstream
-> TT HCM
-> forward RPC to Main HCM
-> Main DB
-> dsyncer
-> TT DB

The goal is narrow: move the call entry point first, keep the write truth in the old path.

To let the old HCM know which data should also flow to TT, add an ownership filter:

agent_id
-> agent_skill_group_rel
-> skill_group
-> access_party
-> TT or Main

Cache the result in Redis. On cache miss, read DB, resolve access party, then write the cache. Work status, agent-skill-group relations, and routing updates can all use this ownership result.

This phase is rollback-friendly. If TT HCM has a problem, traffic can continue through the original main-line path. If the filter has a problem, forwarding can be disabled by config while the old path keeps serving agents.

Version 7 is a migration bridge:

Node / changeProblem it solvesSoftware engineering idea
TT HCM entryUpstreams can move before the write path is fully splitStrangler pattern
Forward RPC to Main HCMOld source of truth stays in charge during phase oneCompatibility adapter
DsyncerTT side can build a local read model while writes remain old-pathData replication
Ownership filterTT and main-line data must be separated by agent ownershipRouting by domain ownership
Redis ownership cacheOwnership checks are hot and repeatedRead-through cache
Config kill switchMigration needs rollback without redeployFeature flag / rollback lever

When can the split become real?

After coexistence is stable, the second phase separates writes, events, MQ, and routing updates.

Version 8: isolate both sides after migration
Figure 9: After the split, TT and Main own separate upstream, HCM, DB, MQ, and routing update paths. Generated by gpt-image-2.

The split includes:

  1. TT HCM stops forwarding RPCs to Main HCM.
  2. The Main-to-TT dsyncer stops.
  3. Main heartbeat and MQ stop consuming TT agent messages.
  4. TT routing consumes only TT status updates.
  5. Main routing consumes only Main status updates.

After that, the old HCM path sheds about 3,000 agents and about 2,000 QPS.

There are two things to prove: the services are separated, and the state facts are closed inside each side.

  • TT agent heartbeat enters only the TT side.
  • TT agent status is committed only by TT HCM.
  • TT status changes go only to TT routing, WFM, and data consumers.
  • The main business line keeps its original path without TT release and cutover risk.

Only then is the split actually done.

Version 8 removes the bridge after ownership is proven:

Node / changeProblem it solvesSoftware engineering idea
Stop forwardingTT writes should no longer depend on main-line availabilityService ownership
Stop dsyncerDual-write/read-model sync should not stay foreverTemporary migration artifact removal
Separate TT / Main MQEvents should stay inside their owning business lineBounded context
Separate routing consumersRouting updates should not cross ownership boundariesConsumer ownership
Closure checksA split is done only when facts and side effects are localInvariant verification

How do we notice when frontend activity stops entering the system?

State calculation depends on input events. Green backend RPC metrics only prove backend services are alive. They do not prove user activity reached Heartbeat.

After one cross-region cutover, agents frequently became abnormal while they were still working in the workbench. The failure was in long-link configuration: keyup and keydown events were not reliably entering Heartbeat.

HCM error rate could not see that part of the input path. We needed input health before state calculation.

Version 9 input health
Figure 10: Before calculating state, the system verifies that user activity enters through Workbench SDK, WS-API, Frontier, Heartbeat intake, and Raw MQ. Generated by gpt-image-2.

Input health needs several dimensions:

MetricWhy it matters
event_rate{event_type, region, channel}Detect drops in click, key, heartbeat, and other activity events
client_lagMeasure delay from client action to server intake
frontier_error_rateDetect long-link channel failures
ws_register_failureDetect module/entity registration failures
short_poll_lagCheck whether short-link compensation keeps up
raw_mq_lagDetect backlog after events enter MQ
synthetic_actionSend scheduled synthetic actions per region to prove the path works
no_action_ratioDetect abnormal candidates concentrated by region, version, or channel

The stronger design feeds input health into the rule layer.

If one region or workbench version has unhealthy input, automatic abnormal rules should degrade: extend the window, pause auto-abnormal, or send reminders only. When input recovers, normal calculation resumes.

That reduces automation during incidents, but it prevents a worse failure: marking working agents abnormal because the collection path broke.

Version 9 adds input observability:

Node / changeProblem it solvesSoftware engineering idea
Input health metricsBackend success does not prove frontend actions arrivedEnd-to-end observability
frontier_error_rate / ws_register_failureLong-link and registration failures need their own signalsLayer-specific telemetry
short_poll_lagDegradation path must be measured tooFallback observability
Synthetic actionPassive metrics may miss a broken path with low trafficActive probing
Rule degradationBad input should reduce automation before it creates bad stateCircuit breaker for business rules

What if HCM is correct and downstream is stale?

Once HCM commits a status, downstream systems still need to receive it. This was the US-region ES incident in another form: HCM DB had the correct offline state, but WFM still displayed the old state.

The binlog-to-HCM-MQ path had about 300k queued messages. The work status handler was not slow. An ES-dependent handler timed out and blocked shared consumption.

The design response is to give state propagation its own lane.

Version 10 handler isolation
Figure 11: Binlog consumers are split by business meaning. A slow search handler no longer blocks work status delivery to WFM and routing. Generated by gpt-image-2.

The target shape treats status change as a first-class event:

HCM transaction
-> status_table
-> status_change_outbox
-> work_status_dispatcher
-> status MQ
-> WFM / Routing

status_change_outbox is written in the same transaction as the status update. The dispatcher only delivers outbox status events to the status MQ.

Search, analytics, and audit can keep consuming binlog or subscribe to their own event streams. They should not share the same blocking point as work status propagation. Slow handlers get their own retry and DLQ.

End-to-end metrics need to follow the same split:

MetricMeaning
status_commit_to_mq_latencyHCM DB commit to status message send
mq_to_wfm_latencyStatus MQ to WFM consumption
handler_lag{handler}Backlog per handler
handler_error_rate{handler}Error rate per handler
downstream_state_ageAge of the state seen by downstream compared with HCM commit

The worst state-system failure is a correct source of truth with stale consumers. Split commit, outbound delivery, and consumption, then measure each segment.

Version 10 isolates propagation:

Node / changeProblem it solvesSoftware engineering idea
status_change_outboxDB commit and outbound event should succeed or fail togetherTransactional outbox
work_status_dispatcherWork status should not wait behind unrelated handlersDedicated worker
Handler-specific retry / DLQSlow consumers need isolated recovery pathsBulkhead isolation
End-to-end latency metricsStale downstream state must be traced to a segmentPipeline observability
Reconciliation directionDownstream state can drift from HCM factsEventual consistency repair

How do we keep hot status reads from taking HCM down?

GetWorkStatus is a hot read in this pipeline. Compute calls it. Upstreams call it. During one OOM incident, downstream error rate spikes aligned with GetWorkStatus traffic spikes.

The state system needs a capacity guard.

Version 11 status read guard
Figure 12: GetWorkStatus gets quota in front, caller isolation inside, and short-TTL cache plus autoscaling signals behind it. Generated by gpt-image-2.

The goal is caller isolation. If one upstream goes bad, its failure should stay inside its budget.

GetWorkStatus needs several defenses:

DefenseJob
Per-caller quotaLimit one abusive upstream first
Bulkhead poolSeparate Compute, WFM, admin, and other callers
Read cacheShort-TTL cache for repeated reads
Stale read budgetAllow slightly stale reads on non-write paths
Circuit breakerFail fast when DB or dependencies are unhealthy
Autoscale signalScale by QPS, heap, GC, and p99 latency

Heartbeat Compute should also reduce pressure. If the same agent reports many events in a short window, Compute does not need to call GetWorkStatus for every event. It can merge by agent in a small window, read a local or Redis snapshot, and still let HCM recheck before final commit.

The path changes from “every caller hits DB-shaped truth” into a read service with budgets, isolation, and degradation.

Version 11 turns a hot read into a protected service:

Node / changeProblem it solvesSoftware engineering idea
Per-caller quotaOne upstream can overwhelm shared state readsFairness and admission control
Bulkhead poolCaller groups should not exhaust each other’s workersBulkhead isolation
Short-TTL cacheRepeated reads should not all hit DBCache-aside
Stale read budgetSome reads can trade freshness for availabilityExplicit consistency budget
Circuit breakerBroken dependencies should fail fastFailure containment
Compute-side mergeMany events for one agent can collapse into fewer readsRequest coalescing

How do query changes ship safely?

One HCM 4.3 rollback exposed a release problem.

Canary looked fine. After ROW, new error logs appeared. Two LEFT JOINs amplified a count SQL path and caused DB timeouts.

Canary alone did not cover this class of failure. Canary traffic is smaller and its data distribution may be kind. Expensive queries need their own cost gate.

Version 12 query release guard
Figure 13: Query changes pass through feature flags, canary, shadow query, EXPLAIN checks, slow SQL monitoring, and rollback gates. Generated by gpt-image-2.

A query rollout can follow this path:

query change
-> feature flag
-> canary
-> shadow query
-> row traffic
-> auto rollback

shadow query does not affect online response. It compares result and cost. The release system watches:

SignalAction
EXPLAIN row estimate over budgetBlock release
Shadow query result mismatchBlock release
p99 query latency over budgetDegrade or rollback
New error logsRollback
DB timeout increaseRollback

Feature flags should be finer than “use new query.” A better split is by filter shape:

simple filter -> old query
join filter -> join query

If a complex filter path breaks, only that path rolls back.

Version 12 adds release gates around data access:

Node / changeProblem it solvesSoftware engineering idea
Feature flagQuery behavior needs runtime rollbackProgressive delivery
CanaryNew query should see a small traffic slice firstControlled exposure
Shadow queryResult and cost can be checked without affecting usersDark launch
EXPLAIN budgetExpensive query plans should be blocked before trafficStatic cost guard
Slow SQL and error budgetsRuntime cost can differ from canary expectationAutomated rollback signal
Filter-level fallbackOne expensive filter should not roll back every query pathGranular kill switch

What remains open?

By this point, the heartbeat design has become a state platform. Intake, compute, commit, propagation, capacity, release, and ownership boundaries can each be observed and degraded separately.

There is still more to build:

QuestionExtension
How do we verify more rules?Build a rule simulator that replays historical events and reports how many agents a new rule would affect
How do we debug one wrong status quickly?Build a per-agent timeline across input events, candidates, zsets, HCM commits, and downstream consumption
Can multi-region ownership split-brain?Drill active_idc switchovers and compare duplicate messages and missing messages
What if downstream stays stale for too long?Add reconciliation jobs that compare WFM/routing state against HCM facts
Can read cache return stale data in dangerous paths?Add versioned cache values and stale-read budgets, while keeping write paths closed by HCM recheck

Final design

The original question was simple: how do we know whether an agent is online?

The final design has two planes.

The data plane turns user activity into state facts:

Agent UI
-> Workbench SDK
-> WS-API
-> Frontier
-> Heartbeat Service
-> Raw MQ
-> Compute
-> Redis zset
-> Exception MQ
-> HCM
-> DB + Outbox
-> Status MQ
-> WFM / Routing / Analytics

The control plane keeps the data plane bounded during failure:

  • The frontend reports activity and heartbeat events.
  • Workbench SDK owns the long-link singleton, reconnect, short-link compensation, and client-side dedupe.
  • Input health proves activity reached Heartbeat.
  • MQ mirror preserves event continuity across regions.
  • Redis zset and dedupe locks collapse repeated candidates.
  • active_idc decides which region can emit exception messages.
  • HCM recheck closes the final status commit.
  • Outbox and handler isolation protect work status propagation.
  • Quota, cache, bulkheads, and circuit breakers protect hot reads.
  • Query gates control risky SQL changes.
  • TT/Main ownership separates agent, skill group, MQ, and routing boundaries.

The design is much larger than last_seen_at, but each layer earns its place. A heartbeat system that drives routing and workforce management is not a ping loop. It is an event intake system, a time-window calculator, a state authority, a propagation pipeline, a multi-region failover path, and a set of release and capacity guardrails around the same fact: what state is this agent in right now?