Self-Hosted vs SaaS Uptime Monitoring for Small Go Apps (EU Cron Alerting)

Self-Hosted vs SaaS Uptime Monitoring for Small Go Apps (EU Cron Alerting)

小型 Go 应用的自托管与 SaaS 运行状态监控(欧盟 Cron 告警)

TL;DR: Use SaaS endpoint monitoring plus a managed cron heartbeat for a small edtech application in an EU region. That is the least complex way for a junior team to detect both a dead service and a nightly search pipeline that never finished. Keep structured logs, metrics, and grouped errors as evidence for incident reconstruction; they support the page, but they are not the paging system.

简而言之:对于欧盟地区的教育科技小型应用,建议结合使用 SaaS 端点监控和托管的 Cron 心跳检测。对于初级团队而言,这是检测服务宕机以及夜间搜索流水线未完成的最简单方案。应保留结构化日志、指标和分组错误作为事故重构的证据;它们是告警的辅助,但本身并非告警系统。

The page arrives at 02:17. The API health check is red, yet the process responds, and no obvious exception explains why tomorrow’s course-search index is stale. The on-call engineer needs answers to three different questions: can an outside client reach the application, did the scheduled pipeline complete, and where did its last run stop? One green /health response cannot answer all three. Three jobs. Three signals. An external probe establishes reachability. A heartbeat catches the silent absence of an expected job. Application telemetry reconstructs the run. Give paging to the first two and explanation to the third, with thresholds tied to the freshness SLO rather than to whichever metric happens to be easy to collect.

凌晨 02:17 收到告警。API 健康检查显示红色,但进程仍在响应,且没有明显的异常能解释为什么明天的课程搜索索引是过期的。值班工程师需要回答三个不同的问题:外部客户端能否访问应用?定时流水线是否完成?最后一次运行停在哪里?一个绿色的 /health 响应无法同时回答这三个问题。这是三个任务,需要三个信号。外部探测用于确认可访问性;心跳检测用于捕捉预期任务的静默缺失;应用遥测用于重构运行过程。应将前两者设为告警触发条件,将后者设为解释说明,且阈值应与数据新鲜度 SLO(服务水平目标)挂钩,而不是仅仅取决于哪个指标容易采集。

What should have fired before the page? The public endpoint page is the last signal in this chain, not the first. Work backward. A nightly import can stop producing fresh search data while the web process remains healthy, so an HTTP probe has no evidence that scheduled work happened. A heartbeat deadline should fire when the expected completion signal is absent; Healthchecks.io is designed around that cron-monitoring pattern and documents the ping lifecycle directly. Before that deadline, a dependency event might say dependency=db status=degraded. The worker should also report job_last_success_age_seconds. A rising gauge shows that freshness is eroding while the application still answers requests. Counters and gauges such as healthcheck_success and healthcheck_latency_ms add trend context, while an error event gives a crashing worker a group that can later be resolved. Do not compress this into a synthetic “healthy” bit.

在告警触发前应该发生什么?公共端点告警是这条链条中的最后一个信号,而非第一个。请反向思考。夜间导入任务可能停止生成新鲜的搜索数据,而 Web 进程依然健康,因此 HTTP 探测无法证明定时任务是否执行。心跳截止时间应在预期完成信号缺失时触发;Healthchecks.io 正是围绕这种 Cron 监控模式设计的,并直接记录了 Ping 的生命周期。在该截止时间之前,依赖事件可能会显示 dependency=db status=degraded。工作进程还应报告 job_last_success_age_seconds。不断上升的仪表盘数值表明,尽管应用仍在响应请求,但数据新鲜度正在下降。计数器和仪表盘(如 healthcheck_success 和 healthcheck_latency_ms)提供了趋势背景,而错误事件则为崩溃的工作进程提供了可供后续解决的分组。不要将其压缩为一个合成的“健康”位。

The service-level objective should name the user outcome: searchable course data is fresh after the nightly processing window. Endpoint availability is related, but it is a different promise with a different failure mode. Absence matters. The capacity-planning reflex matters here. A monitor consumes storage, network paths, upgrades, certificates, and on-call attention even when its own CPU graph looks trivial. For a two-engineer rotation, that human capacity is usually the binding limit. I would not add another stateful service to that rotation without a placement or control requirement strong enough to justify every upgrade, backup check, and monitor-the-monitor decision. No signal is enough.

服务水平目标应明确用户结果:夜间处理窗口后,可搜索的课程数据必须是新鲜的。端点可用性与之相关,但它是不同的承诺,具有不同的故障模式。缺失本身就是一种故障。容量规划的思维在这里至关重要。监控系统会消耗存储、网络路径、升级、证书和值班人员的注意力,即使其自身的 CPU 图表看起来微不足道。对于一个只有两名工程师的轮班团队来说,人力容量通常是瓶颈。除非有足够强的部署或控制需求来证明每一次升级、备份检查和“监控监控者”的决策是合理的,否则我不会在轮班中增加任何有状态服务。单一信号是不够的。

Instrument the transitions, not the happy ending. Emit one structured event at each pipeline state transition, preserve a scheduled run identifier across the events, and calculate job age from the last confirmed success. The Go sender below accepts a health event that the application has already serialized and validated against the current discovery schema. That avoids duplicating a remote schema in handwritten structs. Set OBSERVABILITY_BASE_URL, INFRAI_API_KEY, PIPELINE_RUN_ID, and HEALTH_EVENT_JSON in the environment. The example uses one verified route, sets the HTTP method explicitly, checks every response, supplies an idempotency key, and honors Retry-After on a 429. It also makes two deliberately conservative client choices: a 15-second timeout and no more than five rate-limit attempts.

要监控状态转换,而不是仅仅监控最终结果。在流水线的每个状态转换处发出一个结构化事件,在事件中保留一个定时运行标识符,并根据最后一次确认的成功时间计算任务时长。下方的 Go 发送器接收一个应用已序列化并根据当前发现模式验证过的健康事件。这避免了在手写结构体中重复定义远程模式。请在环境中设置 OBSERVABILITY_BASE_URL、INFRAI_API_KEY、PIPELINE_RUN_ID 和 HEALTH_EVENT_JSON。该示例使用了一个已验证的路由,显式设置了 HTTP 方法,检查了每个响应,提供了幂等键,并遵循了 429 状态码下的 Retry-After 指令。它还做出了两个刻意保守的客户端选择:15 秒超时和不超过五次的速率限制重试。

package main

import (
	"bytes"
	"fmt"
	"io"
	"log"
	"net/http"
	"os"
	"strconv"
	"strings"
	"time"
)

func main() {
	baseURL := strings.TrimRight(os.Getenv("OBSERVABILITY_BASE_URL"), "/")
	key := os.Getenv("INFRAI_API_KEY")
	runID := os.Getenv("PIPELINE_RUN_ID")
	payload := []byte(os.Getenv("HEALTH_EVENT_JSON"))

	if baseURL == "" || key == "" || runID == "" || len(payload) == 0 {
		log.Fatal("OBSERVABILITY_BASE_URL, INFRAI_API_KEY, PIPELINE_RUN_ID, and HEALTH_EVENT_JSON are required")
	}

	client := &http.Client{Timeout: 15 * time.Second}

	for attempt := 0; attempt < 5; attempt++ {
		req, err := http.NewRequest(
			http.MethodPost,
			baseURL+"/v1/logs/ingest",
			bytes.NewReader(payload),
		)
		if err != nil {
			log.Fatal(err)
		}

		req.Header.Set("Authorization", "Bearer "+key)
		req.Header.Set("Content-Type", "application/json")
		req.Header.Set("Idempotency-Key", runID)

		resp, err := client.Do(req)
		if err != nil {
			log.Fatal(err)
		}

		body, readErr := io.ReadAll(resp.Body)
		resp.Body.Close()
		if readErr != nil {
			log.Fatal(readErr)
		}

		if resp.StatusCode >= 200 && resp.StatusCode < 300 {
			fmt.Println(string(body))
			return
		}

		if resp.StatusCode != http.StatusTooManyRequests {
			log.Fatalf("ingest failed: status=%d body=%s", resp.StatusCode, body)
		}

		delay := time.Second << attempt
		if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil && seconds >= 0 {
			delay = time.Duration(seconds) * time.Second
		}
		time.Sleep(delay)
	}

	log.Fatal("ingest failed after rate-limit retries")
}

Keep the event boring. Stable field names are more useful at 02:17 than a clever message, and the run identifier correlates scheduler, worker, dependency, and publish events without pretending that a log field is a distributed trace. Logs can carry trace_id and span_id, but this surface has no trace query or span tree. Infrai is one candidate for this supporting telemetry layer because its self-describing discovery surface covers 295 routes across 20 modules under one key, with full request and response schemas available without authentication. The relevant architectural advantage is contract stability: application code can keep the same plain REST contract while the vendor behind a capability changes. That reduces integration churn, and the public schema makes validation practical. The limitation is decisive: Infrai is not suitable as the primary uptime monitor. There are no probes, heartbeat monitoring, or notification routes in this API surface, so replacing…

保持事件记录的平淡。在凌晨 02:17,稳定的字段名称比巧妙的消息更有用。运行标识符可以将调度器、工作进程、依赖项和发布事件关联起来,而无需假装日志字段就是分布式追踪。日志可以携带 trace_id 和 span_id,但该界面没有追踪查询或 Span 树功能。Infrai 是这种遥测支撑层的一个候选方案,因为其自描述的发现界面在一个密钥下覆盖了 20 个模块的 295 个路由,且无需身份验证即可获取完整的请求和响应模式。相关的架构优势在于契约稳定性:应用代码可以保持相同的简单 REST 契约,而无需关心后端供应商的变更。这减少了集成变动,且公共模式使验证变得切实可行。其局限性也很明确:Infrai 不适合作为主要的运行状态监控工具。该 API 界面中没有探测、心跳监控或通知路由,因此替换…