Production Kubernetes Troubleshooting Lab
本文为原文前 6,000 字符的节选翻译,完整内容请查看原文。
Production Kubernetes Troubleshooting Lab
Observability, Incident Response & Root-Cause Analysis Scenario: You are the DevOps/SRE on-call team for an e-commerce company. Architecture: CUSTOMER │ ▼ order-service │ ▼ Service │ ┌─────────┼─────────┐ ▼ ▼ ▼ Pod-1 Pod-2 Pod-3 │ ▼ Dependencies
可观测性、事件响应与根本原因分析 场景:您是一家电子商务公司的 DevOps/SRE 值班团队。 架构: 客户 │ ▼ 订单服务 │ ▼ 服务 │ ┌─────────┼─────────┐ ▼ ▼ ▼ Pod-1 Pod-2 Pod-3 │ ▼ 依赖项
Today we will deliberately create these incidents: INCIDENT 1 → New deployment → CrashLoopBackOff → Rollback INCIDENT 2 → Production slow → High traffic / CPU → HPA investigation INCIDENT 3 → Running Pods but application unavailable → Service selector INCIDENT 4 → Running 0/1 → Readiness failure INCIDENT 5 → OOMKilled → Memory investigation INCIDENT 6 → Application dependency problem → Logs
今天我们将故意制造以下事件: 事件 1 → 新部署 → CrashLoopBackOff → 回滚 事件 2 → 生产环境缓慢 → 高流量/CPU → HPA 调查 事件 3 → Pod 运行但应用程序不可用 → 服务选择器 事件 4 → 运行 0/1 → 就绪探针失败 事件 5 → OOMKilled → 内存调查 事件 6 → 应用程序依赖问题 → 日志
LAB 0 — Check the Cluster
Everybody starts here.
kubectl cluster-info
Then: kubectl get nodes
Expected:
NAME STATUS ROLES AGE
ip-192-168-10-10.ec2… Ready kubectl get nodes -o wide
Now create our production namespace:
kubectl create namespace production
Check: kubectl get namespace production
实验 0 — 检查集群
每个人都从这里开始。
kubectl cluster-info
然后:kubectl get nodes
预期结果:
名称 状态 角色 年龄
ip-192-168-10-10.ec2… 就绪 kubectl get nodes -o wide
现在创建我们的生产命名空间:
kubectl create namespace production
检查:kubectl get namespace production
LAB 1 — Build Healthy Production
First before troubleshooting production, we need healthy production.
Create: nano production.yaml
Paste:
apiVersion: apps/v1
kind: Deployment
metadata:
name: order-service
namespace: production
spec:
replicas: 3
selector:
matchLabels:
app: order-service
template:
metadata:
labels:
app: order-service
spec:
containers:
- name: order-service
image: nginx:1.27
ports:
- containerPort: 80
resources:
requests:
cpu: “100m”
memory: “64Mi”
limits:
cpu: “500m”
memory: “256Mi”
readinessProbe:
httpGet:
path: /
port: 80
initialDelaySeconds: 5
periodSeconds: 5
livenessProbe:
httpGet:
path: /
port: 80
initialDelaySeconds: 10
periodSeconds: 10
apiVersion: v1 kind: Service metadata: name: order-service namespace: production spec: selector: app: order-service ports:
- port: 80 targetPort: 80 type: ClusterIP
实验 1 — 构建健康的生产环境
在排查生产环境故障之前,我们需要一个健康的生产环境。
创建:nano production.yaml
粘贴:
(此处为 YAML 配置内容,省略重复)
Apply: kubectl apply -f production.yaml
Watch: kubectl get pods -n production -w
Eventually:
NAME READY STATUS
order-service-xxxxxxxxxx-abcde 1/1 Running
order-service-xxxxxxxxxx-fghij 1/1 Running
order-service-xxxxxxxxxx-klmno 1/1 Running
Press: Ctrl+C
应用:kubectl apply -f production.yaml
观察:kubectl get pods -n production -w
最终:
名称 就绪 状态
order-service-xxxxxxxxxx-abcde 1/1 运行中
order-service-xxxxxxxxxx-fghij 1/1 运行中
order-service-xxxxxxxxxx-klmno 1/1 运行中
按:Ctrl+C
Check the Entire Production Environment
kubectl get all -n production
Check labels: kubectl get pods -n production --show-labels
You should see: app=order-service
Check Service: kubectl get svc -n production
Now: kubectl get endpoints order-service -n production
Depending on Kubernetes version, you may also use: kubectl get endpointslice -n production
You should have backend addresses.
检查整个生产环境
kubectl get all -n production
检查标签:kubectl get pods -n production --show-labels
你应该看到:app=order-service
检查服务:kubectl get svc -n production
现在:kubectl get endpoints order-service -n production
根据 Kubernetes 版本,你也可以使用:kubectl get endpointslice -n production
你应该能看到后端地址。
Concept: Service │ │ selector: │ app=order-service │ ▼ Pod Pod Pod app= app= app= order order order -service -service -service Everything matches. Production is healthy.
概念: 服务 │ │ 选择器: │ app=order-service │ ▼ Pod Pod Pod app= app= app= order order order -service -service -service 一切匹配。生产环境是健康的。
Verify the Application
Create a temporary troubleshooting Pod:
kubectl run test-client --image=curlimages/curl --restart=Never -n production -- sleep 3600
Check: kubectl get pod test-client -n production
Then: kubectl exec -n production test-client -- curl -s http://order-service
You should receive the Nginx HTML page.
Now tell students: This is our known-good production baseline. That’s important.
验证应用程序
创建一个临时的排查 Pod:
kubectl run test-client --image=curlimages/curl --restart=Never -n production -- sleep 3600
检查:kubectl get pod test-client -n production
然后:kubectl exec -n production test-client -- curl -s http://order-service
你应该会收到 Nginx HTML 页面。
现在告诉学生:这是我们已知的良好生产基准。这一点很重要。
INCIDENT 1 — New Deployment Breaks Production
STUDENTS ONLY RECEIVE THIS MESSAGE
🚨 INCIDENT: Customers cannot place orders. The Order Service started failing immediately after today’s deployment.
Ask students: What changed?
Answer: New deployment.
First check: kubectl get pods -n production
But currently everything works.
事件 1 — 新部署导致生产环境故障
学生仅收到此消息
🚨 事件:客户无法下单。订单服务在今天部署后立即开始失败。
询问学生:发生了什么变化?
回答:新部署。
首先检查:kubectl get pods -n production
但目前一切正常。
INSTRUCTOR — Break Production
Don’t show students the command.
Run: kubectl set image deployment/order-service order-service=nginx:this-version-does-not-exist -n production
Now students run: kubectl get pods -n production
They should eventually see something similar to:
order-service-old 1/1 Running
order-service-new 0/1 ImagePullBackOff
order-service-new 0/1 ImagePullBackOff
讲师 — 破坏生产环境
不要向学生展示该命令。
运行:kubectl set image deployment/order-service order-service=nginx:this-version-does-not-exist -n production
现在学生运行:kubectl get pods -n production
他们最终应该会看到类似以下内容:
order-service-old 1/1 运行中
order-service-new 0/1 ImagePullBackOff
order-service-new 0/1 ImagePullBackOff
Important teaching point: Because this is a rolling update, old healthy Pods may remain available while new Pods fail. So production may be degraded rather than completely unavailable.
重要的教学点:因为这是滚动更新,旧的健康 Pod 可能保持可用,而新的 Pod 失败。因此,生产环境可能是降级而不是完全不可用。
Student Investigation
First: kubectl rollout status deployment/order-service -n production
It will not complete normally.
Then: kubectl get pods -n production
Pick one failing Pod: kubectl describe pod <NEW-POD-NAME> -n production
Look at Events. Students should find something similar to:
Failed to pull image and: ErrImagePull then: ImagePullBackOff
Now ask: Is ImagePullBackOff the root cause? More precisely, it’s the Kubernetes symptom/state. The useful cause in the Events is that the requested image/tag cannot be pulled.
学生调查
首先:kubectl rollout status deployment/order-service -n production
它不会正常完成。
然后:kubectl get pods -n production
挑选一个失败的 Pod:kubectl describe pod <NEW-POD-NAME> -n production
查看事件。学生应该会发现类似以下内容:
Failed to pull image 以及:ErrImagePull 然后:ImagePullBackOff
现在询问:ImagePullBackOff 是根本原因吗?更准确地说,这是 Kubernetes 的症状/状态。事件中有用的原因是请求的镜像/标签无法拉取。
Check Deployment:
kubectl get deployment order-service -n production -o jsonpath='{.spec.template.spec.containers[0].image}{"\n"}'
It should show: nginx:this-version-does-not-exist
Check Rollout History: kubectl rollout history deployment/order-service -n production
You’ll have revisions.
Now ask: Production worked before the release. The new rollout cannot start. Should we sit in production and experiment with the image? No. Mitigate first.
检查部署:
kubectl get deployment order-service -n production -o jsonpath='{.spec.template.spec.containers[0].image}{"\n"}'
它应该显示:nginx:this-version-does-not-exist
检查发布历史:kubectl rollout history deployment/order-service -n production
你会有修订版本。
现在询问:发布前生产环境是正常的。新的发布无法启动。我们应该在生产环境中尝试修复镜像吗?不。先进行缓解。
kubectl rollout undo deployment/order-service -n production
Then: kubectl rollout status deployment/order-service -n production
Check: kubectl get pods -n production
Eventually:
READY STATUS
1/1 Running
1/1 Running
1/1 Running
Check image: kubectl get deployment order-service -n production -o jsonpath='{.spec.template.spec.containers[0].image}{"\n"}'
Should be back to: nginx:1.27
Verify from customer path: kubectl exec -n production test-client -- curl -s http://order-service
kubectl rollout undo deployment/order-service -n production
然后:kubectl rollout status deployment/order-service -n production
检查:kubectl get pods -n production
最终:
就绪 状态
1/1 运行中
1/1 运行中
1/1 运行中
检查镜像:kubectl get deployment order-service -n production -o jsonpath='{.spec.template.spec.containers[0].image}{"\n"}'
应该恢复为:nginx:1.27
从客户路径验证:kubectl exec -n production test-client -- curl -s http://order-service
Incident #1 conclusion New Deployment ↓ New Pods fail ↓ ImagePullBackOff ↓ Events ↓ Invalid/unavailable image tag ↓ ROLLBACK ↓ Previous known-good release ↓ VERIFY
事件 #1 结论 新部署 ↓ 新 Pod 失败 ↓ ImagePullBackOff ↓ 事件 ↓ 无效/不可用的镜像标签 ↓ 回滚 ↓ 之前已知的良好版本 ↓ 验证
INCIDENT 2 — Production Is Slow
Now we make the lab more professional. Tell students only:
🚨 INCIDENT: Production is slow. Customers report that Order Service responses are taking much longer than normal. There has been no new deployment.
Ask: What’s the cause?
Students should NOT answer: “Traffic.” They should say: We need evidence.
First Establish Baseline
Check: kubectl get pods -n production
All should be: 1/1 Run
事件 2 — 生产环境缓慢
现在我们让实验更专业。仅告诉学生:
🚨 事件:生产环境缓慢。客户报告订单服务响应时间比平时长得多。没有进行新的部署。
询问:原因是什么?
学生不应该回答:“流量”。他们应该说:我们需要证据。
首先建立基准
检查:kubectl get pods -n production
所有都应该是:1/1 运行中