Gaming Cohort Rollbacks: Node.js Cron Heartbeat Health Check for Missed Jobs

Gaming Cohort Rollbacks: Node.js Cron Heartbeat Health Check for Missed Jobs

游戏分组回滚:针对漏跑任务的 Node.js Cron 心跳健康检查

Do not treat a successful retry as proof that a Node.js cron job is healthy. For a gaming experiment, record one durable completion heartbeat per tenant cohort and schedule window, then let an independent watchdog compare those records with explicit deadlines; rollback only the cohorts whose evidence is late or incomplete. Rollback safety is the deciding constraint, because one global green signal can hide a partial rollout that updated the control cohort but skipped a treatment cohort.

不要将成功的重试视为 Node.js 定时任务健康的证明。对于游戏实验,应为每个租户分组(tenant cohort)和调度窗口记录一次持久化的完成心跳,然后让一个独立的监控程序(watchdog)将这些记录与明确的截止时间进行比对;仅对那些证据缺失或延迟的分组执行回滚。回滚安全性是决定性约束,因为一个全局的“健康”信号可能会掩盖部分部署失败的情况——即更新了对照组却跳过了实验组。

TL;DR: Give every expected run a stable key such as experiment_id/cohort/window, write the heartbeat only after the cohort’s durable work commits, and keep retry attempts separate from completion state. A five-minute schedule might use a seven-minute completion deadline, but that number is a capacity decision, not folklore: derive it from the observed upper tail of queue delay plus execution time, add bounded clock-skew allowance, and revisit it when cohort size changes.

简而言之:为每次预期的运行提供一个稳定的键(Key),例如 experiment_id/cohort/window;仅在分组的持久化工作提交后才写入心跳;并将重试尝试与完成状态分开。五分钟的调度周期可能使用七分钟的完成截止时间,但这个数字是基于容量的决策,而非经验主义:应根据观察到的队列延迟上限加上执行时间来推导,并加入有限的时钟偏差余量,且在分组规模变化时重新评估。

How should a Node.js cron health check detect a missed job? Process uptime answers the wrong question. A scheduler can be alive while an invocation never starts; a worker can return success after processing only part of its input; a retry can overlap the next window; and a heartbeat emitted at job start can remain green after the actual write fails. The operational signal must therefore represent the business unit that rollback acts on. Here, that unit is a tenant cohort inside one experiment window.

Node.js 定时任务健康检查应如何检测漏跑的任务?进程的运行时间(uptime)回答的是错误的问题。调度器可能处于活跃状态,但调用从未开始;工作进程可能在仅处理部分输入后就返回成功;重试可能与下一个窗口重叠;而在任务开始时发出的心跳,在实际写入失败后仍可能显示为绿色。因此,运维信号必须代表回滚所作用的业务单元。在这里,该单元是一个实验窗口内的租户分组。

This distinction gets sharp during a staged gaming rollout. Suppose an experiment has control, treatment-a, and treatment-b cohorts on a five-minute cadence. Three expected completion records are required for each window. Two records are not 67% healthy. They mean the window is incomplete, and the safest automated response is to stop advancement for the absent cohort while preserving evidence for the others. No guessing.

在分阶段的游戏部署中,这种区别尤为重要。假设一个实验包含对照组、实验组 A 和实验组 B,周期为五分钟。每个窗口需要三条预期的完成记录。两条记录并不意味着 67% 的健康度。它们意味着窗口未完成,最安全的自动化响应是停止缺失分组的推进,同时保留其他分组的证据。不要猜测。

Retries need their own identity. Keep attempt for diagnosis, but make completion idempotent on the stable run key. Otherwise, attempt two can create a second heartbeat, make counts look complete, or overwrite timing evidence from attempt one. A unique constraint on the run key turns duplicate completion into a harmless no-op; it does not make the underlying cohort mutation idempotent, so that mutation still needs its own transaction or deduplication boundary. Retries lie.

重试需要有自己的标识。保留重试尝试用于诊断,但要使完成状态基于稳定的运行键实现幂等性。否则,第二次尝试可能会创建第二个心跳,使计数看起来已完成,或者覆盖第一次尝试的时间证据。运行键上的唯一性约束将重复的完成操作转化为无害的空操作;但这并不能使底层的分组变更实现幂等,因此该变更仍需要自己的事务或去重边界。重试会撒谎。

Build the completion contract before the alert. The contract needs four times: the scheduled window, the worker’s start, its durable completion, and the watchdog’s observation. Only the first and third decide lateness. Start time is diagnostic, while observation time prevents the checker from pretending it knew something earlier than it did. Store cohort and experiment identifiers as dimensions, not inside a free-form message. A practical state model is small: expected, started, completed, late, and failed. Do not let started satisfy the health check. Likewise, a timeout is not automatically a failure of the mutation; it is an unknown outcome until the durable store is read. That distinction prevents a blind retry from applying the experiment twice.

在告警之前建立完成契约。契约需要四个时间点:调度窗口、工作进程启动、持久化完成以及监控程序的观察时间。只有第一个和第三个时间点决定是否延迟。启动时间用于诊断,而观察时间防止检查程序假装比实际更早获知信息。将分组和实验标识符存储为维度,而不是放在自由格式的消息中。一个实用的状态模型很小:预期(expected)、已启动(started)、已完成(completed)、延迟(late)和失败(failed)。不要让“已启动”状态满足健康检查。同样,超时并不自动意味着变更失败;在读取持久化存储之前,它是一个未知的结果。这种区别可以防止盲目的重试导致实验被应用两次。

The watchdog below is intentionally written in Go and kept outside the Node.js worker. It accepts records through a generic store interface, computes a deadline from the scheduled time, and produces cohort-level decisions. The production adapter should use a transactional datastore with a uniqueness rule over experiment, cohort, and window.

下方的监控程序特意使用 Go 编写,并保持在 Node.js 工作进程之外。它通过通用的存储接口接收记录,根据调度时间计算截止日期,并生成分组级别的决策。生产环境的适配器应使用支持事务的数据库,并对实验、分组和窗口设置唯一性规则。

package watchdog

import (
	"context"
	"fmt"
	"time"
)

type RunKey struct {
	Experiment string
	Cohort     string
	Window     time.Time
}

type Completion struct {
	Key         RunKey
	CompletedAt time.Time
	Attempt     int
}

type Store interface {
	Completion(ctx context.Context, key RunKey) (Completion, bool, error)
}

type Decision struct {
	Key      RunKey
	Rollback bool
	Reason   string
}

func Evaluate(
	ctx context.Context,
	store Store,
	now time.Time,
	grace time.Duration,
	expected []RunKey,
) ([]Decision, error) {
	decisions := make([]Decision, 0, len(expected))
	for _, key := range expected {
		completion, found, err := store.Completion(ctx, key)
		if err != nil {
			return nil, fmt.Errorf("read completion for %s/%s: %w", key.Experiment, key.Cohort, err)
		}

		deadline := key.Window.Add(grace)
		if !found && now.After(deadline) {
			decisions = append(decisions, Decision{key, true, "completion deadline exceeded"})
			continue
		}
		if found && completion.CompletedAt.After(deadline) {
			decisions = append(decisions, Decision{key, true, "completed after deadline"})
			continue
		}
		decisions = append(decisions, Decision{key, false, "within completion contract"})
	}
	return decisions, nil
}

There is a deliberate limitation here: a missing record before its deadline is not healthy or unhealthy yet. It is pending. Alerting early trains responders to ignore noise, while rolling back early can interrupt valid work. On the other side, a very generous grace period protects completion rate by spending rollback time; capacity planning must expose that trade rather than bury it in a timeout constant. The trade is explicit: a seven-minute deadline on a five-minute cadence permits two minutes for queueing, execution variance, and bounded clock skew, but it also means rollback detection cannot be faster than that deadline. If the observed tail no longer fits, shortening the timeout only converts predictable capacity pressure into noisy failures. The honest choices are to add capacity, reduce each window’s cohort work, lengthen the cadence, or accept a slower rollback objective. Record that choice beside the SLO so an operator does not “fix” an alert by widening the grace period during an incident. The deadline wins. The SLO should describe completed cohort windows, not watchdog availability alone. One useful formulation is the proportion of expected cohort windows completed before their deadlines, with missing and late records consuming the error budget. Keep the rollback controller conservative when the evidence store itself is unavailable: freeze rollout progress.

这里有一个刻意的限制:在截止日期前缺失记录既不代表健康也不代表不健康。它处于“待定”状态。过早告警会训练响应者忽略噪音,而过早回滚可能会中断有效的任务。另一方面,过于宽松的宽限期虽然通过牺牲回滚时间保护了完成率,但容量规划必须暴露这种权衡,而不是将其掩盖在超时常量中。这种权衡是明确的:五分钟周期内设置七分钟的截止日期,允许两分钟用于排队、执行偏差和有限的时钟偏移,但也意味着回滚检测不可能快于该截止日期。如果观察到的尾部延迟不再适配,缩短超时只会将可预测的容量压力转化为嘈杂的故障。诚实的做法是增加容量、减少每个窗口的分组工作量、延长周期,或者接受较慢的回滚目标。将此选择记录在 SLO 旁边,这样运维人员就不会在事故期间通过延长宽限期来“修复”告警。截止日期是最终标准。SLO 应描述已完成的分组窗口,而不仅仅是监控程序的可用性。一个有用的公式是:在截止日期前完成的预期分组窗口比例,缺失和延迟的记录将消耗错误预算。当证据存储本身不可用时,保持回滚控制器保守:冻结部署进度。