From $ to # in one unshare call
Technical breakdown of Linux kernel LPE CVEs - nf_tables UAF, Dirty Pipe, StackRot - exploit primitives, attack chains, and the telemetry gap defenders face.
Distribution advisories titled “Several vulnerabilities have been discovered in the Linux kernel” ship on a near-weekly cadence from Debian, Ubuntu, and SUSE. The phrasing is boilerplate. The bugs behind it are not. The pattern under the boilerplate is consistent. A memory-safety defect in a kernel subsystem reachable from an unprivileged context, converted into local privilege escalation. uid 0 from a normal login shell. CVE-2024-1086 is the reference case. Use-after-free in netfilter nf_tables. CVSS v3 7.8. Added to the CISA Known Exploited Vulnerabilities catalog in May 2024. Affected releases span 3.15 through 6.8-rc1, close to a decade of kernels.
The recurring subsystem is netfilter. nf_tables is the packet classification engine behind modern firewalling. It is configured over a netlink socket. That socket is reachable inside a network namespace, and a network namespace can be created by an unprivileged user when CONFIG_USER_NS is enabled and unprivileged clone is permitted. A normal user calls unshare with CLONE_NEWUSER and CLONE_NEWNET, maps themselves to root inside the new namespace, and gains write access to the full nf_tables configuration API. No privilege was granted. The kernel handed a complex, stateful parser to an attacker who controls every byte fed into it.
CVE-2024-1086 lives in the verdict handling path. nft_verdict_init accepted a positive value as a drop error. Later, when a packet hit that rule, nf_hook_slow processed an NF_DROP return carrying a drop error that resembled NF_ACCEPT. The socket buffer backing the packet was freed on the drop path and freed again on the accept path. A double-free of an skb. That is the primitive. The object is freed twice, so two live references believe they own the same slab chunk. Controlled allocation between the two frees places an attacker object on the freed address. The second free then corrupts the slab freelist, and the allocator hands the same page back for a different purpose.
The bug class repeats because the surface is large and stateful. CVE-2023-32233 is a use-after-free in nf_tables anonymous sets, reachable the same way, through user namespaces, producing the same class of arbitrary read and write. The root cause differs. Anonymous sets were handled inconsistently across their API lifetime, so a set could be referenced after it was released within a single transaction. The outcome converges. A dangling kernel object the attacker reclaims with controlled data.
Not every kernel LPE is netfilter. CVE-2022-0847, Dirty Pipe, CVSS v3 7.8, sat in the pipe subsystem. The flags member of a newly allocated pipe_buffer was not initialised in copy_page_to_iter_pipe and push_pipe. A stale PIPE_BUF_FLAG_CAN_MERGE flag, left set from a prior splice, told the kernel it could append into a page-cache page that backed a read-only file. The result is a write primitive against files the user cannot open for writing. Overwrite a page backing /etc/passwd. Overwrite the page cache of a setuid binary. No race window to lose, unlike Dirty COW before it. Linux 5.8 introduced the regression. Fixed in 5.16.11, 5.15.25, and 5.10.102.
CVE-2023-3269, StackRot, CVSS v3 7.8, hit memory management directly. The maple tree replaced the red-black tree for virtual memory area tracking in 6.1. During stack expansion in expand_downwards, a maple tree node was updated without holding the correct lock against concurrent readers, leaving a use-after-free on VMA structures. Affected 6.1 through 6.4. The exploitation is harder, requiring precise control over node reclamation, but the endpoint is the same. Kernel memory the attacker reads and writes.
The exploit path from any of these primitives follows a known shape. Convert the use-after-free or double-free into a page-level reclaim. Spray controlled kernel objects into the freed slab or page so a field the attacker cares about overlaps attacker-controlled bytes. When the vulnerable object lives in a dedicated SLUB cache, the chain adds a cross-cache step. Free an entire slab back to the page allocator, then force the kernel to re-provision that physical page into a different cache holding a more useful victim object. Build a limited read, then a limited write, then widen both into an arbitrary read and arbitrary write across kernel memory. Leak a kernel pointer to defeat KASLR. SMEP and SMAP block execution and data access from user pages, so modern chains stay resident in kernel memory rather than jumping to userland shellcode. The final step is a data-only write. Overwrite modprobe_path so a triggered module autoload executes an attacker path as root. Overwrite core_pattern with a pipe handler so a crash spawns a root helper. Or locate the task’s cred structure and zero the uid and gid fields. No code injection required. The write is the exploit.
This is not theoretical. CVE-2024-1086 has a public analysis and a working exploit demonstrated against multiple distributions with high reliability, and its presence in the CISA KEV catalog means confirmed in-the-wild use. The MITRE mapping is T1068, exploitation for privilege escalation. In a container, the same primitive is T1611, escape to host, because a kernel compromise crosses the namespace boundary that containers depend on. The follow-on is T1548 class abuse of elevation control once root is held. Kernel LPE is the second stage of most real intrusions. Initial access lands code as an unprivileged user or in a constrained container. The kernel bug is what turns that foothold into full host control.
What defenders see is the hard part. The corruption itself is invisible. The double-free, the freelist overwrite, the KASLR leak, the arbitrary write to modprobe_path all happen inside kernel memory. No syscall is emitted for a use-after-free. No EDR hook fires on a slab freelist being corrupted. The telemetry exists only at the edges of the exploit, in the setup before and the behaviour after.
Before the corruption, the setup is loud if logging is tuned for it. The Linux audit framework records clone, unshare, and setns. A rule on unshare with CLONE_NEWUSER from a process that is not a container runtime is a strong signal, and the arg fields record the exact clone flags so the user-plus-network namespace combination can be matched precisely rather than alerting on every container spawn. eBPF-based sensors from Elastic, CrowdStrike, and SentinelOne hook the same syscalls at sys_enter and capture namespace creation with full process ancestry. Falco ships rules for setuid and setgid changes and for namespace manipulation. Sysmon for Linux emits Event ID 1 on process creation, capturing the unshare invocation and its command line. The signal is not the exploit. It is an unprivileged process reaching for user and network namespaces it has no operational reason to touch.
After the corruption, the tell is a privilege state that does not match provenance. The audit login uid, auid, is set at session start and is immutable across the session. An exploit does not change it. A process running with euid 0 while its auid is a normal user id, with no intervening su or sudo execve in the process tree, is privilege escalation that bypassed the sanctioned path. That correlation is the single highest-value detection for kernel LPE. Supporting signals follow the data-only write. A write to /proc/sys/kernel/modprobe. A changed core_pattern that pipes to an executable. A kernel module loaded from a path under /tmp or a user home. Each is anomalous on a production host and each is the visible consequence of a write the sensor could not see land.
The gap is structural. Userland sensors observe syscalls and process events. Kernel memory corruption lives below that boundary. A sensor that watches only process creation and network connections will see the unshare and the eventual root shell and nothing between them. Closing the gap means instrumenting the syscall surface that kernel exploits must traverse for setup, and alerting on the privilege-provenance mismatch that every successful LPE produces, rather than waiting for a kernel-internal event that is never reported.
The residual exposure after patching is the point advisories omit. Each CVE is fixed at a specific boundary. CVE-2024-1086 is resolved after 6.8-rc1 and in the stable backports. Dirty Pipe is closed in the three stable lines named above. But the enabling condition outlives any single patch. Unprivileged user namespaces remain broadly enabled because containers depend on them, and that is the door every netfilter LPE walks through. kernel.unprivileged_userns_clone set to 0 on Debian and Ubuntu, or user.max_user_namespaces set to 0, closes the door where workloads permit. Ubuntu added AppArmor restriction of unprivileged user namespaces as a middle path. Slab hardening reduces reliability. CONFIG_SLAB_FREELIST_HARDENED and CONFIG_SLAB_FREELIST_RANDOM break the freelist assumptions these double-free chains depend on, and CONFIG_INIT_ON_FREE_DEFAULT_ON zeroes freed memory to blunt reclaim sprays. None of these fix the next bug. They raise the cost of exploiting it.
The advisory will read the same next month. Several vulnerabilities have been discovered in the Linux kernel. The subsystem may change from netfilter to the scheduler to a filesystem driver. The shape will not. An unprivileged-reachable memory-safety defect, a reclaim primitive, a data-only write to a root-executed target, and a privilege-provenance mismatch that is the only thing a defender ever gets to see. Operators of systems regulated under the SOCI Act should treat an in-the-wild kernel LPE on a critical asset as a reportable escalation, not a routine patch cycle. Patch to the boundary. Then watch the setup syscalls and the auid mismatch, because the corruption in between will never announce itself.
Keep Reading
Dirty Frag races the refcount
Dirty Frag (CVE-2026-XXXX) is a Linux kernel page migration race yielding root LPE on all major distros. Mechanism, telemetry, and patch boundary.
linux-kernelCVE-2026-31337: Dirty Frag roots every major distro
Technical analysis of CVE-2026-31337 'Dirty Frag': a Linux kernel UAF in IP fragment reassembly giving local root across major distros.
linux-kernelDirty Frag roots every kernel
Technical analysis of CVE-2026-3490 'Dirty Frag' - a page_frag refcount UAF in the Linux kernel enabling local root on stock 5.15-6.8 kernels.
Latest on the Wire
Full wire →- 1936 Electrical Control Room’s Hidden LegacyHacker News
- 40% of health facilities fail mock bird flu outbreak drillArs Technica
- AI Crushes Stratego, Solving Longtime Challenge on a BudgetHacker News
- AI Evangelism Sparks Industry DissatisfactionHacker News
New signal daily · RSS
Stay in the loop
New writing delivered when it's ready. No schedule, no spam.