Production Kubernetes Troubleshooting Lab

本文为原文前 6,000 字符的节选翻译,完整内容请查看原文。

Production Kubernetes Troubleshooting Lab

Observability, Incident Response & Root-Cause Analysis Scenario: You are the DevOps/SRE on-call team for an e-commerce company. Architecture: CUSTOMER │ ▼ order-service │ ▼ Service │ ┌─────────┼─────────┐ ▼ ▼ ▼ Pod-1 Pod-2 Pod-3 │ ▼ Dependencies

可观测性、事件响应与根本原因分析 场景:您是一家电子商务公司的 DevOps/SRE 值班团队。 架构: 客户 │ ▼ 订单服务 │ ▼ 服务 │ ┌─────────┼─────────┐ ▼ ▼ ▼ Pod-1 Pod-2 Pod-3 │ ▼ 依赖项

Today we will deliberately create these incidents: INCIDENT 1 → New deployment → CrashLoopBackOff → Rollback INCIDENT 2 → Production slow → High traffic / CPU → HPA investigation INCIDENT 3 → Running Pods but application unavailable → Service selector INCIDENT 4 → Running 0/1 → Readiness failure INCIDENT 5 → OOMKilled → Memory investigation INCIDENT 6 → Application dependency problem → Logs

今天我们将故意制造以下事件: 事件 1 → 新部署 → CrashLoopBackOff → 回滚 事件 2 → 生产环境缓慢 → 高流量/CPU → HPA 调查 事件 3 → Pod 运行但应用程序不可用 → 服务选择器 事件 4 → 运行 0/1 → 就绪探针失败 事件 5 → OOMKilled → 内存调查 事件 6 → 应用程序依赖问题 → 日志

LAB 0 — Check the Cluster Everybody starts here. kubectl cluster-info Then: kubectl get nodes Expected: NAME STATUS ROLES AGE ip-192-168-10-10.ec2… Ready … ip-192-168-20-20.ec2… Ready … Check: kubectl get nodes -o wide Now create our production namespace: kubectl create namespace production Check: kubectl get namespace production

实验 0 — 检查集群 每个人都从这里开始。 kubectl cluster-info 然后:kubectl get nodes 预期结果: 名称 状态 角色 年龄 ip-192-168-10-10.ec2… 就绪 … ip-192-168-20-20.ec2… 就绪 … 检查:kubectl get nodes -o wide 现在创建我们的生产命名空间: kubectl create namespace production 检查:kubectl get namespace production

LAB 1 — Build Healthy Production First before troubleshooting production, we need healthy production. Create: nano production.yaml Paste: apiVersion: apps/v1 kind: Deployment metadata: name: order-service namespace: production spec: replicas: 3 selector: matchLabels: app: order-service template: metadata: labels: app: order-service spec: containers: - name: order-service image: nginx:1.27 ports: - containerPort: 80 resources: requests: cpu: “100m” memory: “64Mi” limits: cpu: “500m” memory: “256Mi” readinessProbe: httpGet: path: / port: 80 initialDelaySeconds: 5 periodSeconds: 5 livenessProbe: httpGet: path: / port: 80 initialDelaySeconds: 10 periodSeconds: 10

apiVersion: v1 kind: Service metadata: name: order-service namespace: production spec: selector: app: order-service ports:

  • port: 80 targetPort: 80 type: ClusterIP

实验 1 — 构建健康的生产环境 在排查生产环境故障之前,我们需要一个健康的生产环境。 创建:nano production.yaml 粘贴: (此处为 YAML 配置内容,省略重复)

Apply: kubectl apply -f production.yaml Watch: kubectl get pods -n production -w Eventually: NAME READY STATUS order-service-xxxxxxxxxx-abcde 1/1 Running order-service-xxxxxxxxxx-fghij 1/1 Running order-service-xxxxxxxxxx-klmno 1/1 Running Press: Ctrl+C

应用:kubectl apply -f production.yaml 观察:kubectl get pods -n production -w 最终: 名称 就绪 状态 order-service-xxxxxxxxxx-abcde 1/1 运行中 order-service-xxxxxxxxxx-fghij 1/1 运行中 order-service-xxxxxxxxxx-klmno 1/1 运行中 按:Ctrl+C

Check the Entire Production Environment kubectl get all -n production Check labels: kubectl get pods -n production --show-labels You should see: app=order-service Check Service: kubectl get svc -n production Now: kubectl get endpoints order-service -n production Depending on Kubernetes version, you may also use: kubectl get endpointslice -n production You should have backend addresses.

检查整个生产环境 kubectl get all -n production 检查标签:kubectl get pods -n production --show-labels 你应该看到:app=order-service 检查服务:kubectl get svc -n production 现在:kubectl get endpoints order-service -n production 根据 Kubernetes 版本,你也可以使用:kubectl get endpointslice -n production 你应该能看到后端地址。

Concept: Service │ │ selector: │ app=order-service │ ▼ Pod Pod Pod app= app= app= order order order -service -service -service Everything matches. Production is healthy.

概念: 服务 │ │ 选择器: │ app=order-service │ ▼ Pod Pod Pod app= app= app= order order order -service -service -service 一切匹配。生产环境是健康的。

Verify the Application Create a temporary troubleshooting Pod: kubectl run test-client --image=curlimages/curl --restart=Never -n production -- sleep 3600 Check: kubectl get pod test-client -n production Then: kubectl exec -n production test-client -- curl -s http://order-service You should receive the Nginx HTML page. Now tell students: This is our known-good production baseline. That’s important.

验证应用程序 创建一个临时的排查 Pod: kubectl run test-client --image=curlimages/curl --restart=Never -n production -- sleep 3600 检查:kubectl get pod test-client -n production 然后:kubectl exec -n production test-client -- curl -s http://order-service 你应该会收到 Nginx HTML 页面。 现在告诉学生:这是我们已知的良好生产基准。这一点很重要。

INCIDENT 1 — New Deployment Breaks Production STUDENTS ONLY RECEIVE THIS MESSAGE 🚨 INCIDENT: Customers cannot place orders. The Order Service started failing immediately after today’s deployment. Ask students: What changed? Answer: New deployment. First check: kubectl get pods -n production But currently everything works.

事件 1 — 新部署导致生产环境故障 学生仅收到此消息 🚨 事件:客户无法下单。订单服务在今天部署后立即开始失败。 询问学生:发生了什么变化? 回答:新部署。 首先检查:kubectl get pods -n production 但目前一切正常。

INSTRUCTOR — Break Production Don’t show students the command. Run: kubectl set image deployment/order-service order-service=nginx:this-version-does-not-exist -n production Now students run: kubectl get pods -n production They should eventually see something similar to: order-service-old 1/1 Running order-service-new 0/1 ImagePullBackOff order-service-new 0/1 ImagePullBackOff

讲师 — 破坏生产环境 不要向学生展示该命令。 运行:kubectl set image deployment/order-service order-service=nginx:this-version-does-not-exist -n production 现在学生运行:kubectl get pods -n production 他们最终应该会看到类似以下内容: order-service-old 1/1 运行中 order-service-new 0/1 ImagePullBackOff order-service-new 0/1 ImagePullBackOff

Important teaching point: Because this is a rolling update, old healthy Pods may remain available while new Pods fail. So production may be degraded rather than completely unavailable.

重要的教学点:因为这是滚动更新,旧的健康 Pod 可能保持可用,而新的 Pod 失败。因此,生产环境可能是降级而不是完全不可用。

Student Investigation First: kubectl rollout status deployment/order-service -n production It will not complete normally. Then: kubectl get pods -n production Pick one failing Pod: kubectl describe pod <NEW-POD-NAME> -n production Look at Events. Students should find something similar to: Failed to pull image and: ErrImagePull then: ImagePullBackOff Now ask: Is ImagePullBackOff the root cause? More precisely, it’s the Kubernetes symptom/state. The useful cause in the Events is that the requested image/tag cannot be pulled.

学生调查 首先:kubectl rollout status deployment/order-service -n production 它不会正常完成。 然后:kubectl get pods -n production 挑选一个失败的 Pod:kubectl describe pod <NEW-POD-NAME> -n production 查看事件。学生应该会发现类似以下内容: Failed to pull image 以及:ErrImagePull 然后:ImagePullBackOff 现在询问:ImagePullBackOff 是根本原因吗?更准确地说,这是 Kubernetes 的症状/状态。事件中有用的原因是请求的镜像/标签无法拉取。

Check Deployment: kubectl get deployment order-service -n production -o jsonpath='{.spec.template.spec.containers[0].image}{"\n"}' It should show: nginx:this-version-does-not-exist Check Rollout History: kubectl rollout history deployment/order-service -n production You’ll have revisions. Now ask: Production worked before the release. The new rollout cannot start. Should we sit in production and experiment with the image? No. Mitigate first.

检查部署: kubectl get deployment order-service -n production -o jsonpath='{.spec.template.spec.containers[0].image}{"\n"}' 它应该显示:nginx:this-version-does-not-exist 检查发布历史:kubectl rollout history deployment/order-service -n production 你会有修订版本。 现在询问:发布前生产环境是正常的。新的发布无法启动。我们应该在生产环境中尝试修复镜像吗?不。先进行缓解。

kubectl rollout undo deployment/order-service -n production Then: kubectl rollout status deployment/order-service -n production Check: kubectl get pods -n production Eventually: READY STATUS 1/1 Running 1/1 Running 1/1 Running Check image: kubectl get deployment order-service -n production -o jsonpath='{.spec.template.spec.containers[0].image}{"\n"}' Should be back to: nginx:1.27 Verify from customer path: kubectl exec -n production test-client -- curl -s http://order-service

kubectl rollout undo deployment/order-service -n production 然后:kubectl rollout status deployment/order-service -n production 检查:kubectl get pods -n production 最终: 就绪 状态 1/1 运行中 1/1 运行中 1/1 运行中 检查镜像:kubectl get deployment order-service -n production -o jsonpath='{.spec.template.spec.containers[0].image}{"\n"}' 应该恢复为:nginx:1.27 从客户路径验证:kubectl exec -n production test-client -- curl -s http://order-service

Incident #1 conclusion New Deployment ↓ New Pods fail ↓ ImagePullBackOff ↓ Events ↓ Invalid/unavailable image tag ↓ ROLLBACK ↓ Previous known-good release ↓ VERIFY

事件 #1 结论 新部署 ↓ 新 Pod 失败 ↓ ImagePullBackOff ↓ 事件 ↓ 无效/不可用的镜像标签 ↓ 回滚 ↓ 之前已知的良好版本 ↓ 验证

INCIDENT 2 — Production Is Slow Now we make the lab more professional. Tell students only: 🚨 INCIDENT: Production is slow. Customers report that Order Service responses are taking much longer than normal. There has been no new deployment. Ask: What’s the cause? Students should NOT answer: “Traffic.” They should say: We need evidence. First Establish Baseline Check: kubectl get pods -n production All should be: 1/1 Run

事件 2 — 生产环境缓慢 现在我们让实验更专业。仅告诉学生: 🚨 事件:生产环境缓慢。客户报告订单服务响应时间比平时长得多。没有进行新的部署。 询问:原因是什么? 学生不应该回答:“流量”。他们应该说:我们需要证据。 首先建立基准 检查:kubectl get pods -n production 所有都应该是:1/1 运行中