field notes · falcosecurity/plugins · issue #1500

The Anatomy of an Undefined Symbol

 
no code changed, only linkage, verified on Debian 11, Ubuntu 24.04 and aarch64
updated. The fix merged, and the CI guard that followed it got a review that found my verification was only half done. Section 8 now covers what the guard missed and why, which turned out to be the same lesson as the bug itself.
A chain of linked blue blocks reaching toward a missing dashed block, while an unused green block sits off to the side

A security sensor that works everywhere modern silently refuses to start on half of enterprise Linux. Traced from a red dashboard to a two-line fix, now submitted upstream. Along the way: one cheap experiment, four lines of readelf, and a linker you can step through yourself. Budget ~15 minutes.

Keep one question in mind through every section: “how can a binary that links successfully still be broken, and where does the breakage hide until it’s on a customer’s machine?”

1. A dashboard full of red

~1 min · read

I run a home lab that boots real VMs with real kernels, 3.10 through 6.8, and installs Falco (falcosecurity/falco), the CNCF runtime security sensor, on each one. Not containers: containers share the host kernel, and when the thing you’re testing is kernel-dependent, a container is just the host wearing a costume. The lab measures what a customer would actually get. Does the package install? Does the service start? Which eBPF driver does it pick? Does it actually detect anything?

One night the legacy tier came back all red. Five distros (Debian 11, Rocky 8, Ubuntu 20.04, Ubuntu 18.04, CentOS 7) and Falco running on zero of them.

The obvious read: old kernels, unsupported, done. That read was wrong on every single row. This post is about the most interesting one.

2. Read the journal, not the summary

~2 min · read

On a kept-alive Debian 11 VM, journalctl showed Falco crash-looping. systemd restarted it every 15 seconds, forever. The actual error, identical every cycle:

Runtime error: cannot load plugin /usr/share/falco/plugins/libcontainer.so: can't load plugin dynamic library: /usr/share/falco/plugins/libcontainer.so: undefined symbol: __res_search. Exiting.

Two things worth noticing before touching anything:

• This happens before any engine opens. Falco can run on modern eBPF, legacy eBPF, or a kernel module, and this failure kills all three identically, because plugin loading precedes engine selection. Which is exactly why it looks like “this kernel is unsupported” on a matrix.
• The default ruleset requires this plugin. So the failure mode isn’t “container metadata unavailable.” It’s the sensor failing to start at all, with stock configuration, on Debian 11, RHEL/Rocky/Alma 8, and Ubuntu 20.04. Distros the package installs on without complaint.

__res_search is a glibc resolver symbol, DNS-lookup plumbing. Why would a container-metadata plugin need DNS? Because it embeds a Go library (built with -buildmode=c-archive) that talks to Docker/containerd/CRI-O, and Go’s net package, since Go 1.20, calls res_search from libresolv through cgo.

So the plugin genuinely uses the symbol. The question is why the loader can’t find it, and why only on some machines.

3. The one-variable experiment

~2 min · read

Before reading a single line of build code, there’s a cheap experiment that tests the whole hypothesis. If the problem is that a library containing __res_search never gets loaded, then forcing that library into the process should flip the result, with zero other changes:

Lab · run the experiment yourself same binary, one variable
// choose a command to run

Same binary. Same config. One environment variable. Broken to working.

What you just saw: the symbol exists on this system, in libresolv.so.2, sitting right there in /lib. The dynamic loader simply never had a reason to load it. Which means the bug isn’t in the code. It’s in the binary’s declared dependencies.

(That LD_PRELOAD line is also a legitimate stopgap: one systemd drop-in and a stock broken install runs. But a workaround is a tourniquet, not a fix.)

4. Reading the binary’s dependency list

~2 min · read

Every ELF shared object carries a list of libraries it needs, the DT_NEEDED entries, which is what the dynamic loader walks at dlopen time. Here’s the shipped plugin’s:

Lab · inspect the binary readelf, both worlds

There’s the whole bug in four lines: the binary says “I need __res_search” and simultaneously says “I depend on nothing that provides it.” libresolv isn’t in the list.

So why does this same binary work fine on Ubuntu 24.04?

Because glibc 2.34 merged libresolv into libc. On any glibc newer than 2.34, libc.so.6 itself exports __res_search as a compatibility symbol, and libc is always loaded. The missing dependency is invisibly papered over on every modern machine. On glibc 2.28 through 2.33 (Debian 11, RHEL 8, Ubuntu 20.04), the symbol lives only in libresolv.so.2, which nobody asked for. dlopen fails.

Lab · where does __res_search live?drag through glibc history

This is why the bug survived in shipped releases for a year, and why two earlier bug reports went stale without a fix: everyone who could reproduce it was on old glibc; everyone who could fix it was on new. Their CI even builds on Debian bullseye specifically to keep old-glibc compatibility, and this one missing link flag defeats the entire effort.

5. The mechanism: how the linker walks the line

~4 min · step the linker yourself

The plugin’s CMake, it turns out, knows about the Go resolver requirement. It even cites the Go release notes. And then it handles it for macOS only:

if(APPLE) find_library(RESOLV resolv REQUIRED) ... set(WORKER_DEP ${SECURITY_FRAMEWORK} ${RESOLV} ${CORE}) endif()

No Linux branch. Half of a correct fix, shipped for the platform where the bug barely matters.

So the fix is obvious: add set(WORKER_DEP resolv) for Linux. I did exactly that, rebuilt, watched the build succeed with exit code zero, and checked the result: nothing changed. The fix compiled, linked, and silently didn’t take.

To see why, you have to walk the link line the way GNU ld does: a single pass, left to right, carrying a running list of “symbols someone needs that nobody has provided yet.” Try it yourself:

Lab · step the linker
[plugin .o files]-lresolvlibworker.a
needed symbols
(empty)
DT_NEEDED (output .so)
(none yet)
// press "step" to start the walk

What you just saw (in the default order): the linker reaches -lresolv while the needed-list is still empty, discards it, and never goes back. Then libworker.a announces __res_search, too late. And because unresolved symbols are allowed by design in shared libraries (they’re assumed to resolve at load time), the link exits zero and the broken artifact ships. Now press swap order and run it again. Same two ingredients, opposite outcome. Then flip the mode to Apple’s ld64, where order doesn’t matter, and you’ll see why upstream never noticed: the ordering bug was unobservable on the platform the code was written for.

The actual fix is therefore two lines, and the second one is the load-bearing one:

# 1. declare the dependency on Linux (mirroring the APPLE branch) elseif(CMAKE_HOST_SYSTEM_NAME STREQUAL "Linux") set(WORKER_DEP resolv) # 2. and link it AFTER the archive that needs it -target_link_libraries(container PRIVATE ... ${WORKER_DEP} ${WORKER_LIB}) +target_link_libraries(container PRIVATE ... ${WORKER_LIB} ${WORKER_DEP})

Needs before provides. Objects and archives first, the libraries that satisfy them after. The one-line verification that separates believing from knowing:

$ readelf -d libcontainer.so | grep resolv (NEEDED) Shared library: [libresolv.so.2]

6. Trust, but verify, on three axes

~2 min · read

A fix that changes linkage metadata touches every platform, so the verification matrix has to say more than “works on my repro box”:

• Debian 11 (glibc 2.31): was a crash loop on boot. Now: stock config runs under systemd, default rules load, detections fire, events are enriched.
• Ubuntu 24.04 (glibc 2.39): already worked. After the patch: identical behavior, no regression.
• Linux aarch64: builds clean, same NEEDED entry, symbol correctly versioned @GLIBC_2.17.
• macOS / Windows: untouched by construction. The APPLE branch is unchanged and WORKER_DEP is never set on Windows.

The strongest argument is structural: the patch changes zero lines of code. The plugin always called res_search; the fix only writes down the dependency it always had. The kernel-matrix lab that found the bug became the test rig that proved the fix. Every row above is a real VM, not a container.

One honest boundary surfaced during verification: on Ubuntu 18.04 (glibc 2.27) the plugin fails for a different reason, version ‘GLIBC_2.28’ not found. That’s the plugin’s build-baseline floor, and no link flag lowers it. Knowing exactly where a fix stops working is part of the fix.

7. Shipping it upstream

~1 min · read

The failure had been reported twice before, as falco#3719 and falco#3728, and both went stale and closed unfixed. Fair enough: a symptom report without a mechanism is easy to lose. What I filed instead:

• falcosecurity/plugins#1500: the mechanism (missing -lresolv, so no DT_NEEDED), the readelf receipts, the LD_PRELOAD experiment, the affected-distro list, and the stopgap.
• falcosecurity/plugins#1501: the two-line CMake fix, DCO-signed, with a reviewer note about the link-order trap, because a fix that can silently fail deserves a warning label. (Merged 3 September 2026.)

Eleven lines of diff. About four hundred lines of evidence. That ratio is the job.

8. The test that could not have existed

~2 min · read

The fix merged, and the maintainer added two notes that were more interesting than the merge. The first explained why this bug reached a release at all:

leogr, on the pull request

“None of our CI jobs load libcontainer.so on a glibc < 2.34 host (build-linux builds on bullseye but never dlopens the result, and falco-tests runs inside falcosecurity/falco:master-debian, i.e. Debian 12 with glibc 2.36), so CI could not catch this.”

Read that carefully, because it is a precise description of a blind spot rather than an apology. Two jobs touch this library. One builds it on Debian 11, glibc 2.31, old enough to have the bug, and never loads it. The other loads it, inside Debian 12, glibc 2.36, new enough that the bug cannot appear. Each job holds one half of the condition and neither holds both.

That is the same shape as the bug itself. Section 5 argued that glibc 2.34 papers over the missing dependency wherever modern tooling runs, so the fault only surfaces where the fix authors are not standing. Their CI was standing in exactly the same place.

Making it un-reintroducible

A merged fix removes the bug. It does not stop the next Go or CMake change from dropping -lresolv again, silently, on a build machine where nothing will notice. So the follow-up matters more than the fix: plugins#1513 adds two checks to build-linux, the job that was already running on bullseye and already had everything needed.

readelf -d libcontainer.so | grep -q 'NEEDED.*libresolv\.so\.2'

That asserts the dependency is recorded. It is fast and its failure names the cause. But it only ever catches this one symptom, so the second check does the thing CI was never doing: compiles a small harness and dlopens the built library with RTLD_NOW, on glibc 2.31, in the job that just produced it.

RTLD_NOW rather than RTLD_LAZY is the whole point. Lazy binding defers function symbol resolution until first call, so a library missing res_search would load cleanly and fail later, somewhere less obvious. Binding everything at load time is what turns a latent fault into a red build.

Before opening it I checked the guard rather than assuming. Built a probe library calling res_search two ways inside debian:bullseye:

glibc: ldd (Debian GLIBC 2.31-13+deb11u14) 2.31 --- built with -lresolv --- readelf guard: PASS dlopen guard : PASS --- built without -lresolv --- readelf guard: FAIL dlopen guard : FAIL

The second case is this entire post reproduced in four lines of shell. A shared object links happily without -lresolv, because undefined symbols are permitted at link time, and then refuses to load anywhere res_search still lives in libresolv.

I wrote “proved” in the first version of this section. It was not proved, and the next section is about how I found that out.

Where that verification was hollow

The review came back changes requested, and it opened with something I had not done:

leogr, on the pull request

“I ran the two new steps against the real libcontainer-amd64 artifact built from main inside debian:bullseye. The readelf step passes, but the dlopen step fails with undefined symbol: pthread_mutex_trylock. As is, this would turn the next plugins/container/** PR red.”

The probe was linked with -ldl alone. The real plugin embeds a Go runtime that calls pthread_* and dl*, but declares no dependency on the libraries those live in. Its whole DT_NEEDED list is libresolv.so.2, libc.so.6 and the loader. On glibc < 2.34 those functions sit in libpthread and libdl, exactly as res_search sat in libresolv.

It works in production because Falco already links both through libsinsp, so a plugin loaded into that process resolves them from the global scope. My probe was a bare program that linked almost nothing. It handed the plugin an emptier world than it ever ships into, and demanded self-sufficiency the library has never needed.

So the guard would have failed every healthy build. Not the bug it was written to catch. Every other one.

What the table above was actually worth

Both rows are true, and together they establish half of what matters. They show the guard rejects a broken library. They say nothing about whether it accepts a working one, because the only library I ever pointed it at was one I had written to be broken.

For a gate, that is the cheaper half. A check that misses a regression costs you the regression. A check that fires on healthy builds blocks everyone until somebody deletes it, and then you have neither the check nor the regression caught.

The fixture was the trap. My test object was two lines of C. The real one is 37MB with a language runtime inside it. A fixture you build yourself contains only what you thought to put in it, which means it can only ever test the failure you already imagined.

The fix links the probe the way Falco links, so it loads the plugin under the conditions the plugin actually ships into:

-gcc -o /tmp/dlopen_check /tmp/dlopen_check.c -ldl +gcc -o /tmp/dlopen_check /tmp/dlopen_check.c -Wl,--no-as-needed -lpthread -ldl

--no-as-needed is load bearing. Linkers drop libraries the program does not itself call, and the probe never calls pthread_create, so a plain -lpthread would be discarded and nothing would change.

Then the verification I should have run the first time, against the real artifact from the main branch and against the regression, using the workflow’s steps extracted from the YAML rather than my approximation of them:

real libcontainer.so, old probe readelf PASS dlopen FAIL <- false positive real libcontainer.so, fixed probe readelf PASS dlopen PASS .so missing -lresolv, fixed probe readelf FAIL dlopen FAIL <- #1500 still caught

The middle row is the one that was missing. The bottom row is the one that matters after loosening a check, because a guard relaxed until it stops crying wolf can quietly stop catching wolves.

One more thing fell out of the review. The workflow only triggered on plugins/container/**, so the steps I added never ran on the pull request that added them. The green checks came from unrelated workflows, and I had read them as evidence. Adding the workflow to its own path filter fixed that, and build-linux now runs on changes to itself, which is how the corrected probe came to be tested on amd64 and arm64 in their CI rather than only in my container.

The lesson is not about linkers. When a bug survives a test suite, the useful question is not “why did nobody write this test” but “what would this test have had to run on?” The answer here was an environment the project builds on constantly and never executes in. Finding a bug is worth something. Removing the conditions that let it hide is worth more.

The question turned out to cut both ways. Their test suite never ran on an old glibc with the library actually loaded. My test of that suite never ran on a real library. Same question, one level up, and I did not think to ask it of my own work until somebody else did. That is the part I would keep if I could only keep one thing from this.

9. Check yourself

~1 min · answer before revealing
1 · A shared library links with an unresolved symbol and exit code 0. Bug or feature?
Feature, by design, for shared objects: their symbols are assumed to resolve at load time. Which is exactly what makes it a great place for bugs to hide. For an executable, the same situation is a hard link error you’d catch immediately.
2 · gcc ... -lfoo bar.a where bar.a needs symbols from libfoo. What happens on GNU ld, and on Apple’s ld64?
GNU ld: foo is discarded before bar.a announces its needs. Broken output, clean exit. ld64: fine, it resolves across the whole input set regardless of order. Same command line, different linkers, opposite results.
3 · Your fix adds a library to the link line and the build passes. What single command tells you whether the fix actually took?
readelf -d out.so | grep NEEDED. Trust the dynamic table, not the exit code. The build system will happily produce a byte-for-byte equally broken binary with a green checkmark on it.
4 · Why did this bug reproduce for users but never for maintainers?
glibc 2.34 merged libresolv into libc, so every modern build and test machine papers over the missing dependency automatically. The bug only exists where the fix authors aren’t standing.
5 · A CI guard passes every test its author runs and is still wrong. What was not tested?
That it accepts a healthy build. Testing that a check catches the bad case is the instinct, because that is what the check is for. The other direction is the expensive one: a check that misses a regression costs you the regression, while a check that fires on good builds blocks everyone until someone deletes it, and then you have neither. The trap here was the fixture. A two line shared object written to be broken cannot exhibit the properties of a 37MB library with a language runtime in it.
The bug hunt happened in my kernel compatibility matrix, a KVM lab that boots real kernels from 3.10 to 6.8 and characterises a security sensor on each: which driver it lands on, what it costs at idle, and whether it actually detects. The full findings write-up, including four other bugs this investigation shook loose, lives in the repo’s FINDINGS.md.