Calling a function in C without naming it
本文为原文前 6,000 字符的节选翻译,完整内容请查看原文。
Calling a function in C without naming it 05 October 2026
在 C 语言中调用函数而不直接命名它 2026年10月5日
My school has a controlled remote code execution environment to automatically grade the code we submit. The system checks execution, verifies proper formatting and surely has other checks not shown to the user (e.g. a code plagiarism detection tool1). The interface reports results mismatches and compilation errors.
我的学校有一个受控的远程代码执行环境,用于自动评分我们提交的代码。该系统会检查执行情况、验证格式是否正确,并且肯定还有其他未向用户展示的检查(例如代码查重工具)。界面会报告结果不匹配和编译错误。
A friend and I started to search for a way to have a single solution that could solve every exercise. The first step in that direction was to have a primitive that could give us arbitrary shell code execution. We assumed there would be an sh binary on the grading VM such that we could simply call execve. So… are we done? You guessed it, no.
我和一位朋友开始寻找一种能够解决所有练习的通用方案。实现这一目标的第一步是获得一个能够执行任意 shell 代码的原语。我们假设评分虚拟机上存在一个 sh 二进制文件,这样我们就可以简单地调用 execve。那么……我们成功了吗?你猜对了,没有。
Remember when I told that the grading system had other sanity checks? Well, one of them is to check that you only use functions allowed in the subject of the exercise. And execve is never one of them, so it fails the submission before your code even gets to run. From what we experienced, the system is a little more involved than a simple grep of the source code. My guess is that it uses clang to pre-process and parse the code and then does its checks against the resulting AST.
还记得我提到评分系统有其他完整性检查吗?其中之一就是检查你是否只使用了练习题目中允许的函数。而 execve 从来不在允许列表中,因此在你的代码运行之前,提交就会失败。根据我们的经验,该系统不仅仅是对源代码进行简单的 grep 检查。我猜测它使用 clang 对代码进行预处理和解析,然后针对生成的抽象语法树(AST)进行检查。
This is where we get to the core of the problem. We need to call execve but we are not allowed to name the symbol. So how can we do it? Zooming back a bit, we know that our C code essentially gets compiled to a binary and calling a function boils down to a jump to some place in memory. Could we not retrieve the address of execve in memory, hardcode it in our program and call it directly?
这就是问题的核心所在。我们需要调用 execve,但不允许直接使用该符号。那么我们该怎么做呢?退一步讲,我们知道 C 代码最终会被编译成二进制文件,而调用函数归根结底就是跳转到内存中的某个位置。我们难道不能获取 execve 在内存中的地址,将其硬编码到程序中并直接调用吗?
In modern laptop software, memory addresses are often execution-dependent. This happens because modern kernels implement Address Space Layout Randomization (ASLR) which randomizes addresses at which your binary, its libraries, heap and stack are placed. Nowadays, C code compiles using the -pie flag by default, this creates a position-independent executable and means that your binary can indeed take advantage of ASLR.
在现代笔记本电脑软件中,内存地址通常与执行过程相关。这是因为现代内核实现了地址空间布局随机化(ASLR),它会随机化二进制文件、库、堆和栈的存放地址。如今,C 代码默认使用 -pie 标志进行编译,这会创建位置无关的可执行文件,意味着你的二进制文件确实可以利用 ASLR。
We could use the -no-pie C compiler flag to disable this behavior and force fixed addresses for functions in our binary. This doesn’t help because we most often do not control the build process in our case. Second, the execve function is not directly part of our executable but dynamically loaded, so it still is at a random place. Static addresses won’t work.
我们可以使用 -no-pie C 编译器标志来禁用此行为,并强制二进制文件中的函数使用固定地址。但这没有帮助,因为在我们的案例中,我们通常无法控制构建过程。其次,execve 函数并非直接包含在我们的可执行文件中,而是动态加载的,因此它仍然处于随机位置。静态地址是行不通的。
But ASLR only acts on the mapping of the segments, their content is left intact. Inside a segment, functions are still ordered in the same way. If we manage to know the static offset between the memory addresses of two functions fA and fB, obtaining the address of fA allows us to construct the address of fB and vice-versa. The offset varies from a machine to another because of a different version of the C library, or maybe a different architecture, or maybe for some other reason.
但 ASLR 只作用于段的映射,其内容保持不变。在同一个段内,函数的排列顺序依然相同。如果我们能获知两个函数 fA 和 fB 内存地址之间的静态偏移量,那么获取 fA 的地址就能让我们推导出 fB 的地址,反之亦然。由于 C 库版本不同、架构不同或其他原因,偏移量在不同机器上会有所差异。
There are multiple ways to obtain the fixed offset between two functions on the target machine. One of them is to just leak the offset in an exercise where you can name symbols of both fA and fB. Our target function is still execve which is part of libc. We assume we always have access to printf. We just have to leak the offset between the two functions on the VM.
在目标机器上获取两个函数之间固定偏移量的方法有很多。其中一种是在一个允许命名 fA 和 fB 符号的练习中直接泄露该偏移量。我们的目标函数仍然是属于 libc 的 execve。我们假设始终可以访问 printf。我们只需要在虚拟机上泄露这两个函数之间的偏移量即可。
Let’s test this locally first: printf(“%td”, (ptrdiff_t)((size_t)execve - (size_t)printf)); Running the program a couple of times gives us different results. Didn’t we just say that the offset was fixed? Let’s open the binary to investigate. $ objdump —dynamic-syms ./offset | grep -E “\s(execve|printf)” 0000000000000000 DF UND 0000000000000000 (GLIBC_2.2.5) execve 000000000004dd64 w DF .text 0000000000000005 Base printf
让我们先在本地测试一下:printf(“%td”, (ptrdiff_t)((size_t)execve - (size_t)printf)); 多次运行该程序,结果却各不相同。我们刚才不是说偏移量是固定的吗?让我们打开二进制文件调查一下。$ objdump —dynamic-syms ./offset | grep -E “\s(execve|printf)” 0000000000000000 DF UND 0000000000000000 (GLIBC_2.2.5) execve 000000000004dd64 w DF .text 0000000000000005 Base printf
objdump is useful for any kind of object file analysis. The —dynamic-syms option gives information about each symbol that will be loaded at runtime. The sixth column gives us a bit of information about where the symbol comes from. execve indeed comes from glibc, which is the implementation of the C standard library available on my laptop. But printf doesn’t seem to come from glibc. Actually, the raw output of objdump has a bunch of _interceptor* symbols which gives a clue on what’s going on here.
objdump 对于任何类型的目标文件分析都很有用。—dynamic-syms 选项提供了有关运行时将加载的每个符号的信息。第六列提供了一些关于符号来源的信息。execve 确实来自 glibc,这是我笔记本电脑上可用的 C 标准库实现。但 printf 似乎并非来自 glibc。实际上,objdump 的原始输出中有一堆 _interceptor* 符号,这为我们了解这里发生了什么提供了线索。
Since the beginning, I have been compiling all my code with ASan (a.k.a -fsanitize=address) to catch memory errors early. ASan works by replacing a bunch of code, including memory reads and write, by a bunch of slow function calls that check the correctness of each action. It also replaces a bunch of functions in the C standard library like malloc or printf. By doing so, these functions do not live in the glibc object anymore as opposed to functions that aren’t replaced, in our case execve.
从一开始,我就一直使用 ASan(即 -fsanitize=address)编译所有代码,以便尽早发现内存错误。ASan 的工作原理是将大量代码(包括内存读写)替换为一系列缓慢的函数调用,以检查每个操作的正确性。它还替换了 C 标准库中的许多函数,如 malloc 或 printf。这样一来,这些函数就不再位于 glibc 对象中,而那些未被替换的函数(如本例中的 execve)则依然保留在原处。
The automatic grader system also happens to check for memory errors and compiles with ASan. We could try to use a different symbol than printf that lives in the same segment as execve, but we still need a symbol that is allowed most of the time. Most of these symbols live in the ASan replaced segment, so let’s commit to it. Maybe execve is not the only way.
自动评分系统恰好也会检查内存错误并使用 ASan 进行编译。我们可以尝试使用除 printf 之外、且与 execve 位于同一段的其他符号,但我们仍然需要一个大多数情况下都被允许使用的符号。这些符号大多位于被 ASan 替换的段中,所以我们还是接受这一点吧。也许 execve 并不是唯一的途径。
Looking for other useful symbols in this segment, mmap catches my eye. Why is mmap special? Because of the W^X (write-xor-execute) security policy, there is no section of memory that is writable and executable at the same time. Thus we cannot just write raw x86 instructions in a buffer and jump to it as if it was a function nor can we overwrite an existing function with our own code. But mmap solves this problem because we can just tell it to give us a page that has write and execute permissions.
在寻找该段中其他有用的符号时,mmap 引起了我的注意。为什么 mmap 很特别?由于 W^X(写异或执行)安全策略,内存中不存在既可写又可执行的区域。因此,我们不能简单地将原始 x86 指令写入缓冲区并像调用函数一样跳转到它,也不能用我们自己的代码覆盖现有函数。但 mmap 解决了这个问题,因为我们可以要求它提供一个具有写和执行权限的页面。
But let’s not get ahead of ourselves. We can leak the offset between printf and mmap on the VM via another exercise where we can name both symbols. To use the offset we could use a bunch of casts but the code we hand to the grading system has to compile with -pedantic and this disallows cast from, to and between function pointers types. Because function pointers are still pointers in the end, we can try to modify it without the compiler noticing. I started to modify memory based
但我们不要操之过急。我们可以通过另一个允许命名 printf 和 mmap 符号的练习,在虚拟机上泄露它们之间的偏移量。为了使用这个偏移量,我们可以使用一系列类型转换,但我们提交给评分系统的代码必须通过 -pedantic 编译,而这禁止了函数指针类型之间、以及函数指针与其他类型之间的转换。由于函数指针最终仍然是指针,我们可以尝试在不被编译器察觉的情况下修改它。我开始基于内存进行修改。