The Anatomy of an Undefined Symbol

A security sensor that works everywhere modern silently refuses to start on half of enterprise Linux. Traced from a red dashboard to a two-line fix, now submitted upstream. Along the way: one cheap experiment, four lines of readelf, and a linker you can step through yourself. Budget ~15 minutes.
1. A dashboard full of red
I run a home lab that boots real VMs with real kernels, 3.10 through 6.8, and installs Falco (falcosecurity/falco), the CNCF runtime security sensor, on each one. Not containers: containers share the host kernel, and when the thing you’re testing is kernel-dependent, a container is just the host wearing a costume. The lab measures what a customer would actually get. Does the package install? Does the service start? Which eBPF driver does it pick? Does it actually detect anything?
One night the legacy tier came back all red. Five distros (Debian 11, Rocky 8, Ubuntu 20.04, Ubuntu 18.04, CentOS 7) and Falco running on zero of them.
The obvious read: old kernels, unsupported, done. That read was wrong on every single row. This post is about the most interesting one.
2. Read the journal, not the summary
On a kept-alive Debian 11 VM, journalctl showed Falco crash-looping. systemd restarted it every 15 seconds, forever. The actual error, identical every cycle:
Two things worth noticing before touching anything:
• This happens before any engine opens. Falco can run on modern eBPF, legacy eBPF, or a kernel module, and this failure kills all three identically, because plugin loading precedes engine selection. Which is exactly why it looks like “this kernel is unsupported” on a matrix.
• The default ruleset requires this plugin. So the failure mode isn’t “container metadata unavailable.” It’s the sensor failing to start at all, with stock configuration, on Debian 11, RHEL/Rocky/Alma 8, and Ubuntu 20.04. Distros the package installs on without complaint.
__res_search is a glibc resolver symbol, DNS-lookup plumbing. Why would a container-metadata plugin need DNS? Because it embeds a Go library (built with -buildmode=c-archive) that talks to Docker/containerd/CRI-O, and Go’s net package, since Go 1.20, calls res_search from libresolv through cgo.
So the plugin genuinely uses the symbol. The question is why the loader can’t find it, and why only on some machines.
3. The one-variable experiment
Before reading a single line of build code, there’s a cheap experiment that tests the whole hypothesis. If the problem is that a library containing __res_search never gets loaded, then forcing that library into the process should flip the result, with zero other changes:
Same binary. Same config. One environment variable. Broken to working.
What you just saw: the symbol exists on this system, in libresolv.so.2, sitting right there in /lib. The dynamic loader simply never had a reason to load it. Which means the bug isn’t in the code. It’s in the binary’s declared dependencies.
(That LD_PRELOAD line is also a legitimate stopgap: one systemd drop-in and a stock broken install runs. But a workaround is a tourniquet, not a fix.)
4. Reading the binary’s dependency list
Every ELF shared object carries a list of libraries it needs, the DT_NEEDED entries, which is what the dynamic loader walks at dlopen time. Here’s the shipped plugin’s:
There’s the whole bug in four lines: the binary says “I need __res_search” and simultaneously says “I depend on nothing that provides it.” libresolv isn’t in the list.
So why does this same binary work fine on Ubuntu 24.04?
Because glibc 2.34 merged libresolv into libc. On any glibc newer than 2.34, libc.so.6 itself exports __res_search as a compatibility symbol, and libc is always loaded. The missing dependency is invisibly papered over on every modern machine. On glibc 2.28 through 2.33 (Debian 11, RHEL 8, Ubuntu 20.04), the symbol lives only in libresolv.so.2, which nobody asked for. dlopen fails.
This is why the bug survived in shipped releases for a year, and why two earlier bug reports went stale without a fix: everyone who could reproduce it was on old glibc; everyone who could fix it was on new. Their CI even builds on Debian bullseye specifically to keep old-glibc compatibility, and this one missing link flag defeats the entire effort.
5. The mechanism: how the linker walks the line
The plugin’s CMake, it turns out, knows about the Go resolver requirement. It even cites the Go release notes. And then it handles it for macOS only:
No Linux branch. Half of a correct fix, shipped for the platform where the bug barely matters.
So the fix is obvious: add set(WORKER_DEP resolv) for Linux. I did exactly that, rebuilt, watched the build succeed with exit code zero, and checked the result: nothing changed. The fix compiled, linked, and silently didn’t take.
To see why, you have to walk the link line the way GNU ld does: a single pass, left to right, carrying a running list of “symbols someone needs that nobody has provided yet.” Try it yourself:
[plugin .o files]-lresolvlibworker.aWhat you just saw (in the default order): the linker reaches -lresolv while the needed-list is still empty, discards it, and never goes back. Then libworker.a announces __res_search, too late. And because unresolved symbols are allowed by design in shared libraries (they’re assumed to resolve at load time), the link exits zero and the broken artifact ships. Now press swap order and run it again. Same two ingredients, opposite outcome. Then flip the mode to Apple’s ld64, where order doesn’t matter, and you’ll see why upstream never noticed: the ordering bug was unobservable on the platform the code was written for.
The actual fix is therefore two lines, and the second one is the load-bearing one:
Needs before provides. Objects and archives first, the libraries that satisfy them after. The one-line verification that separates believing from knowing:
6. Trust, but verify, on three axes
A fix that changes linkage metadata touches every platform, so the verification matrix has to say more than “works on my repro box”:
• Debian 11 (glibc 2.31): was a crash loop on boot. Now: stock config runs under systemd, default rules load, detections fire, events are enriched.
• Ubuntu 24.04 (glibc 2.39): already worked. After the patch: identical behavior, no regression.
• Linux aarch64: builds clean, same NEEDED entry, symbol correctly versioned @GLIBC_2.17.
• macOS / Windows: untouched by construction. The APPLE branch is unchanged and WORKER_DEP is never set on Windows.
The strongest argument is structural: the patch changes zero lines of code. The plugin always called res_search; the fix only writes down the dependency it always had. The kernel-matrix lab that found the bug became the test rig that proved the fix. Every row above is a real VM, not a container.
One honest boundary surfaced during verification: on Ubuntu 18.04 (glibc 2.27) the plugin fails for a different reason, version ‘GLIBC_2.28’ not found. That’s the plugin’s build-baseline floor, and no link flag lowers it. Knowing exactly where a fix stops working is part of the fix.
7. Shipping it upstream
The failure had been reported twice before, as falco#3719 and falco#3728, and both went stale and closed unfixed. Fair enough: a symptom report without a mechanism is easy to lose. What I filed instead:
• falcosecurity/plugins#1500: the mechanism (missing -lresolv, so no DT_NEEDED), the readelf receipts, the LD_PRELOAD experiment, the affected-distro list, and the stopgap.
• falcosecurity/plugins#1501: the two-line CMake fix, DCO-signed, with a reviewer note about the link-order trap, because a fix that can silently fail deserves a warning label. (Merged 3 September 2026.)
Eleven lines of diff. About four hundred lines of evidence. That ratio is the job.
8. The test that could not have existed
The fix merged, and the maintainer added two notes that were more interesting than the merge. The first explained why this bug reached a release at all:
“None of our CI jobs load libcontainer.so on a glibc < 2.34 host (build-linux builds on bullseye but never dlopens the result, and falco-tests runs inside falcosecurity/falco:master-debian, i.e. Debian 12 with glibc 2.36), so CI could not catch this.”
Read that carefully, because it is a precise description of a blind spot rather than an apology. Two jobs touch this library. One builds it on Debian 11, glibc 2.31, old enough to have the bug, and never loads it. The other loads it, inside Debian 12, glibc 2.36, new enough that the bug cannot appear. Each job holds one half of the condition and neither holds both.
That is the same shape as the bug itself. Section 5 argued that glibc 2.34 papers over the missing dependency wherever modern tooling runs, so the fault only surfaces where the fix authors are not standing. Their CI was standing in exactly the same place.
Making it un-reintroducible
A merged fix removes the bug. It does not stop the next Go or CMake change from dropping -lresolv again, silently, on a build machine where nothing will notice. So the follow-up matters more than the fix: plugins#1513 adds two checks to build-linux, the job that was already running on bullseye and already had everything needed.
readelf -d libcontainer.so | grep -q 'NEEDED.*libresolv\.so\.2'
That asserts the dependency is recorded. It is fast and its failure names the cause. But it only ever catches this one symptom, so the second check does the thing CI was never doing: compiles a small harness and dlopens the built library with RTLD_NOW, on glibc 2.31, in the job that just produced it.
RTLD_NOW rather than RTLD_LAZY is the whole point. Lazy binding defers function symbol resolution until first call, so a library missing res_search would load cleanly and fail later, somewhere less obvious. Binding everything at load time is what turns a latent fault into a red build.
Before opening it I checked the guard rather than assuming. Built a probe library calling res_search two ways inside debian:bullseye:
The second case is this entire post reproduced in four lines of shell. A shared object links happily without -lresolv, because undefined symbols are permitted at link time, and then refuses to load anywhere res_search still lives in libresolv.
I wrote “proved” in the first version of this section. It was not proved, and the next section is about how I found that out.
Where that verification was hollow
The review came back changes requested, and it opened with something I had not done:
“I ran the two new steps against the real libcontainer-amd64 artifact built from main inside debian:bullseye. The readelf step passes, but the dlopen step fails with undefined symbol: pthread_mutex_trylock. As is, this would turn the next plugins/container/** PR red.”
The probe was linked with -ldl alone. The real plugin embeds a Go runtime that calls pthread_* and dl*, but declares no dependency on the libraries those live in. Its whole DT_NEEDED list is libresolv.so.2, libc.so.6 and the loader. On glibc < 2.34 those functions sit in libpthread and libdl, exactly as res_search sat in libresolv.
It works in production because Falco already links both through libsinsp, so a plugin loaded into that process resolves them from the global scope. My probe was a bare program that linked almost nothing. It handed the plugin an emptier world than it ever ships into, and demanded self-sufficiency the library has never needed.
So the guard would have failed every healthy build. Not the bug it was written to catch. Every other one.
Both rows are true, and together they establish half of what matters. They show the guard rejects a broken library. They say nothing about whether it accepts a working one, because the only library I ever pointed it at was one I had written to be broken.
For a gate, that is the cheaper half. A check that misses a regression costs you the regression. A check that fires on healthy builds blocks everyone until somebody deletes it, and then you have neither the check nor the regression caught.
The fixture was the trap. My test object was two lines of C. The real one is 37MB with a language runtime inside it. A fixture you build yourself contains only what you thought to put in it, which means it can only ever test the failure you already imagined.
The fix links the probe the way Falco links, so it loads the plugin under the conditions the plugin actually ships into:
--no-as-needed is load bearing. Linkers drop libraries the program does not itself call, and the probe never calls pthread_create, so a plain -lpthread would be discarded and nothing would change.
Then the verification I should have run the first time, against the real artifact from the main branch and against the regression, using the workflow’s steps extracted from the YAML rather than my approximation of them:
The middle row is the one that was missing. The bottom row is the one that matters after loosening a check, because a guard relaxed until it stops crying wolf can quietly stop catching wolves.
One more thing fell out of the review. The workflow only triggered on plugins/container/**, so the steps I added never ran on the pull request that added them. The green checks came from unrelated workflows, and I had read them as evidence. Adding the workflow to its own path filter fixed that, and build-linux now runs on changes to itself, which is how the corrected probe came to be tested on amd64 and arm64 in their CI rather than only in my container.
The lesson is not about linkers. When a bug survives a test suite, the useful question is not “why did nobody write this test” but “what would this test have had to run on?” The answer here was an environment the project builds on constantly and never executes in. Finding a bug is worth something. Removing the conditions that let it hide is worth more.
The question turned out to cut both ways. Their test suite never ran on an old glibc with the library actually loaded. My test of that suite never ran on a real library. Same question, one level up, and I did not think to ask it of my own work until somebody else did. That is the part I would keep if I could only keep one thing from this.
9. Check yourself
1 · A shared library links with an unresolved symbol and exit code 0. Bug or feature?
2 · gcc ... -lfoo bar.a where bar.a needs symbols from libfoo. What happens on GNU ld, and on Apple’s ld64?
foo is discarded before bar.a announces its needs. Broken output, clean exit. ld64: fine, it resolves across the whole input set regardless of order. Same command line, different linkers, opposite results.3 · Your fix adds a library to the link line and the build passes. What single command tells you whether the fix actually took?
readelf -d out.so | grep NEEDED. Trust the dynamic table, not the exit code. The build system will happily produce a byte-for-byte equally broken binary with a green checkmark on it.