The forgetful CPU (Linux on M4)

The forgetful CPU (Linux on M4)

健忘的 CPU(在 M4 上运行 Linux)

This blog post goes into quite some detail about how I first booted Linux on my M4 Mac mini. I encourage you to look up terms and concepts you don’t know, since I cannot explain all the background in this post ;) Before saying anything further, I need to thank the entire Asahi Linux team for all their prior work and help during the journey. Please consider donating to the Asahi Open Collective if you want to see more mainline Linux on Apple Silicon work!

这篇博文详细介绍了我是如何首次在 M4 Mac mini 上启动 Linux 的。我鼓励你去查阅那些你不熟悉的术语和概念,因为我无法在这篇文章中解释所有的背景知识 ;) 在继续之前,我需要感谢整个 Asahi Linux 团队在这一过程中所做的所有前期工作和提供的帮助。如果你希望看到更多关于 Apple Silicon 主线 Linux 的工作,请考虑向 Asahi Open Collective 捐款!

The Beginnings

起步

In November 2024, I bought an M4 Mac mini, gambling that it would be similar to the M1-M3 Apple Silicon machines and could be quickly supported in Asahi Linux. While the M4 was sitting on my desk for several months, more details started surfacing about this SoC. It turned out to be more difficult, since the M4 machines are the first generation of Apple Silicon to mandate SPTM (Secure Page Table Monitor), which provides hardening against vulnerabilities in the XNU kernel of macOS.

2024 年 11 月,我买了一台 M4 Mac mini,赌它会和 M1-M3 Apple Silicon 机器相似,并能很快得到 Asahi Linux 的支持。当这台 M4 在我的桌子上放了几个月后,关于这款 SoC 的更多细节开始浮出水面。事实证明这更困难,因为 M4 机器是第一代强制要求使用 SPTM(安全页表监视器)的 Apple Silicon,它为 macOS 的 XNU 内核提供了针对漏洞的加固。

In previous generations, the Linux bringup was largely based on MMIO traces captured using the m1n1 hypervisor, allowing the analysis of interactions between the original macOS drivers and the hardware. With SPTM, major changes to m1n1 are required to get macOS running under the hypervisor, and these changes certainly go beyond what I could come up with as a newbie in this space. Still, it wasn’t entirely a lost cause. In parallel to the hypervisor work, I started my first attempts to boot Linux on the M4. This meant disabling strict boot security, installing m1n1 as a custom boot object via macOS recovery, and obtaining a serial console to inspect the logs.

在之前的几代产品中,Linux 的启动主要基于使用 m1n1 虚拟机管理程序捕获的 MMIO 跟踪,这允许分析原始 macOS 驱动程序与硬件之间的交互。有了 SPTM,需要对 m1n1 进行重大修改才能让 macOS 在虚拟机管理程序下运行,这些修改显然超出了我这个领域新手的能力范围。不过,这并非完全无望。在进行虚拟机管理程序工作的同时,我开始了在 M4 上启动 Linux 的首次尝试。这意味着需要禁用严格的启动安全性,通过 macOS 恢复模式安装 m1n1 作为自定义启动对象,并获取串行控制台以检查日志。

Locked registers

被锁定的寄存器

Initially, m1n1 was only able to start in BRINGUP mode and immediately crashed while attempting to initialize GXF. GXF functionality turned out to be disabled/locked in raw boot mode on these M4+ SoCs, so making its initialization conditional and skipping it on these machines was the correct thing to do. Besides the disabled GXF feature, there was also the RVBAR (Reset Vector Base Address Register): a place in memory (one per CPU core) that determines where the core starts executing when powered on. The m1n1 code writes its entrypoint address to the RVBAR for each core when booting the kernel or chainloading another m1n1. Writing to it also led to a crash on M4, but long story short, it already contained the correct value so this write needed to be skipped as well.

最初,m1n1 只能在 BRINGUP 模式下启动,并且在尝试初始化 GXF 时立即崩溃。事实证明,在这些 M4+ SoC 上,GXF 功能在原始启动模式下是被禁用或锁定的,因此将它的初始化设为条件执行并在这些机器上跳过它是正确的做法。除了被禁用的 GXF 功能外,还有 RVBAR(重置向量基地址寄存器):这是内存中的一个位置(每个 CPU 核心一个),决定了核心在通电时从何处开始执行。当引导内核或链式加载另一个 m1n1 时,m1n1 代码会将入口点地址写入每个核心的 RVBAR。在 M4 上写入该寄存器也会导致崩溃,但长话短说,它已经包含了正确的值,因此这个写入操作也需要被跳过。

debug_putc

调试输出

At this point, a long time went by without me touching the Mac mini. At the Chaos Communication Congress at the end of 2025, I found some new motivation. Finally, I hacked together a very minimal device tree containing only the CPU cores and AIC interrupt controller and loaded the Linux kernel (using m1n1’s linux.py) with the earlycon parameter, but I did not see any output after “Vectoring to next stage”.

此时,我很久没有再碰这台 Mac mini 了。在 2025 年底的混沌通信大会(Chaos Communication Congress)上,我找到了一些新的动力。最终,我拼凑出了一个非常精简的设备树,仅包含 CPU 核心和 AIC 中断控制器,并使用 earlycon 参数加载了 Linux 内核(使用 m1n1 的 linux.py),但在“Vectoring to next stage”之后我没有看到任何输出。

Since I wasn’t getting any useful output from the kernel, I could only take wild guesses at what was going wrong, right? I decided to try the brute-force method: good old println-debugging. I took the debug_putc assembly routine from m1n1 and adjusted it to print a single ‘a’ character. I inserted this into the Linux kernel code very early in boot, and sure enough, I got an ‘a’ after “Vectoring to next stage”! Essentially, I bisected the Linux boot code and landed at the MMU init code (which is still very early, still in the assembly code in arch/arm64/kernel/head.S).

由于我无法从内核获得任何有用的输出,我只能对出错的原因进行盲目猜测,对吧?我决定尝试暴力破解法:老派的 println 调试。我从 m1n1 中提取了 debug_putc 汇编例程,并将其调整为打印单个字符“a”。我将其插入到 Linux 内核启动非常早期的代码中,果然,在“Vectoring to next stage”之后我得到了一个“a”!本质上,我对 Linux 启动代码进行了二分查找,最终定位到了 MMU 初始化代码(这仍然非常早期,还在 arch/arm64/kernel/head.S 的汇编代码中)。

Was the MMU initialization crashing the CPU somehow? Not quite: the UART is accessed using memory-mapped I/O. Once the MMU is enabled, all memory accesses are directed to virtual addresses, which are mapped to a corresponding physical addresses through the page-tables. While m1n1 creates mappings to expose the MMIO address space at identical virtual addresses, Linux does not do this, meaning we end up accessing unmapped space instead of the UART once the MMU is enabled. I modified the initial pagetables to add this 1:1 mapping for the MMIO space, and now my debug_putc worked much further into the boot process.

是 MMU 初始化导致 CPU 崩溃了吗?不完全是:UART 是通过内存映射 I/O (MMIO) 访问的。一旦启用了 MMU,所有的内存访问都会指向虚拟地址,这些地址通过页表映射到相应的物理地址。虽然 m1n1 创建了映射以在相同的虚拟地址处暴露 MMIO 地址空间,但 Linux 并没有这样做,这意味着一旦启用 MMU,我们最终访问的是未映射的空间,而不是 UART。我修改了初始页表以添加此 MMIO 空间的 1:1 映射,现在我的 debug_putc 在启动过程中能工作得更远了。

Bisecting again, the print now worked up to somewhere in the interrupt controller initialization! I narrowed it down to a write to the implementation-specific CPU register SYS_IMP_APL_VM_TMR_FIQ_ENA_EL2, which triggered the new crash. After I commented out this write, the kernel booted to a shell. Let’s goooo! SYS_IMP_APL_VM_TMR_FIQ_ENA_EL2, which is related to virtualization, has since been unlocked in new iBoot versions, so commenting out the write is no longer necessary.

再次进行二分查找,打印功能现在可以工作到中断控制器初始化的某个地方了!我将其缩小到对特定于实现的 CPU 寄存器 SYS_IMP_APL_VM_TMR_FIQ_ENA_EL2 的写入,这触发了新的崩溃。在我注释掉这个写入操作后,内核成功启动到了 shell。太棒了!SYS_IMP_APL_VM_TMR_FIQ_ENA_EL2 与虚拟化相关,此后在新的 iBoot 版本中已被解锁,因此不再需要注释掉该写入操作。

Secondary cores and WFI

次级核心与 WFI

For now, m1n1 didn’t start the secondary cores because smp_start_offset was missing. I tried the offset used for base M1 - M3 and was able to start the secondary cores. Once I attempted loading Linux again, I arrived at another mysterious crash. Previous Apple Silicon CPUs already had known quirks regarding the WFI instruction. Depending on the state of the chicken bit (a bit in a register, that allows the vendor to disable some CPU optimization or feature) called ARM64_REG_CYC_OVRD_ok2pwrdn_force_mask in the XNU OSS code, the WFI instruction causes the CPU registers x0-x31 to be zeroed on these previous generations.

目前,m1n1 没有启动次级核心,因为缺少 smp_start_offset。我尝试了基础 M1-M3 使用的偏移量,并成功启动了次级核心。当我再次尝试加载 Linux 时,又遇到了另一次神秘的崩溃。之前的 Apple Silicon CPU 在 WFI 指令方面已经有已知的怪癖。根据 XNU OSS 代码中名为 ARM64_REG_CYC_OVRD_ok2pwrdn_force_mask 的“chicken bit”(寄存器中的一位,允许供应商禁用某些 CPU 优化或功能)的状态,WFI 指令会导致这些前几代产品的 CPU 寄存器 x0-x31 被清零。