Event-Driven ML Pipeline Orchestration for Manufacturing: An AWS Industry Experience
Title: Event-Driven ML Pipeline Orchestration for Manufacturing: An AWS Industry Experience
标题:面向制造业的事件驱动型机器学习流水线编排:AWS 行业实践经验
Abstract: We present an industry experience report on three years of operating an event-driven cloud infrastructure for continuous machine learning training in automotive manufacturing. Our system orchestrates GPU-accelerated training of product-specialized model pairs, a physics prediction model and a reinforcement-learning control policy, across multiple plants, coordinating long-running GPU workloads triggered by manufacturing events.
摘要:我们提交了一份关于在汽车制造业中运行事件驱动型云基础设施以进行持续机器学习训练的三年行业经验报告。我们的系统在多个工厂间编排了针对特定产品模型对(物理预测模型和强化学习控制策略)的 GPU 加速训练,并协调由制造事件触发的长时间运行的 GPU 工作负载。
The architecture combines Amazon ECS with EC2 GPU capacity providers, SQS-based messaging with dead-letter queues, and an admission-controlled Lambda dispatcher that enforces cluster concurrency limits. A Conductor orchestrator on ECS Fargate initiates dependency-aware retraining chains on a weekly schedule. The entire infrastructure is codified in modular Terraform with multi-account separation.
该架构结合了 Amazon ECS 与 EC2 GPU 容量提供商、基于 SQS 的消息传递(含死信队列)以及强制执行集群并发限制的准入控制 Lambda 调度器。运行在 ECS Fargate 上的 Conductor 编排器按周计划启动具有依赖感知能力的再训练链。整个基础设施均采用模块化 Terraform 编写,并实现了多账户隔离。
From 40000+ production training jobs we report a 72-78% cost reduction versus always-on GPU infrastructure. A discrete-event simulation confirms that admission control is necessary (naive dispatch loses 65% of jobs) and that queue-draining matches AWS Step Functions latency while eliminating per-job startup overhead. We provide lessons learned and release the simulator and Terraform module skeletons as open-source artifacts.
通过 40,000 多项生产训练任务,我们报告称,与始终在线的 GPU 基础设施相比,成本降低了 72-78%。离散事件模拟证实了准入控制的必要性(简单调度会丢失 65% 的任务),并表明队列排空机制在匹配 AWS Step Functions 延迟的同时,消除了每个任务的启动开销。我们总结了经验教训,并将模拟器和 Terraform 模块框架作为开源工件发布。