summaryrefslogtreecommitdiff
path: root/fs
AgeCommit message (Collapse)AuthorFilesLines
2026-06-29ovl: handle idmapped mounts in ovl_set_acl()Christian Brauner1-3/+3
The two checks ovl_set_acl() performs on the overlay inode itself - inode_owner_or_capable() and the setgid-stripping test - were done with &nop_mnt_idmap. On an idmapped overlay mount this compares the overlay inode's id space directly against the caller's: it denies the rightful owner (as seen through the mount idmap) the right to set an ACL, and evaluates the setgid-drop decision in the wrong id space, which can mis-set the mode on the upper inode. Use the struct mnt_idmap passed in by the VFS for both. Fold the open-coded "caller not in group and not CAP_FSETID privileged" test into in_group_or_capable() with i_gid_into_vfsgid(), the same idmap-aware helpers used by setattr_should_drop_sgid(). The subsequent internal forced setgid-kill via ovl_setattr() stays on &nop_mnt_idmap: it carries only ATTR_KILL_SGID with no uid/gid to translate and is overlayfs' own mode change, not a user-driven operation through the mount. No functional change until FS_ALLOW_IDMAP is set on ovl_fs_type. Link: https://patch.msgid.link/20260615-work-idmapped-overlayfs-v1-5-7381632aa402@kernel.org Reviewed-by: Amir Goldstein <amir73il@gmail.com> Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
2026-06-29ovl: handle idmapped mounts in ovl_getattr()Christian Brauner1-0/+8
ovl_getattr() fetches attributes from the real (upper or lower) path via vfs_getattr(), so the returned stat->uid/stat->gid are already mapped through the real layer idmap but not through the overlay mount idmap. Unlike the generic path, the VFS does not re-apply the accessing mount's idmap after ->getattr returns - generic_fillattr() does it, but overlayfs bypasses it - so overlayfs has to do it itself. Map stat->uid/stat->gid through the overlay mount idmap before returning, mirroring generic_fillattr(). The owner reported through an idmapped overlay mount is thus the overlay-final id translated by the mount idmap. No functional change until FS_ALLOW_IDMAP is set on ovl_fs_type; until then make_vfsuid()/make_vfsgid() are the identity for &nop_mnt_idmap. Link: https://patch.msgid.link/20260615-work-idmapped-overlayfs-v1-4-7381632aa402@kernel.org Reviewed-by: Amir Goldstein <amir73il@gmail.com> Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
2026-06-29ovl: handle idmapped mounts in ovl_setattr()Christian Brauner1-1/+9
Pass the mount's struct mnt_idmap to setattr_prepare() so that the permission checks for a chown/chmod performed through an idmapped overlay mount are evaluated in the mount's id space. The ownership requested in @attr is expressed relative to the overlay mount idmap. Before forwarding the change to the upper layer via ovl_do_notify_change() - whose notify_change() applies the upper layer idmap in turn - rebase ia_vfsuid/ia_vfsgid into the overlay's own id space, i.e. the same space as the overlay inode's i_{u,g}id established by ovl_copyattr(). Without this rebase the upper layer would interpret the caller's mount-relative id as an upper-relative one and store the wrong owner on disk, or reject it with -EOVERFLOW. from_vfsuid() returns INVALID_UID for an id that the overlay mount idmap does not map; that invalid id is carried faithfully into the forwarded iattr and rejected by the upper notify_change() via vfsuid_has_fsmapping(), so no bogus owner can be written. No functional change until FS_ALLOW_IDMAP is set on ovl_fs_type; until then the overlay mount idmap is &nop_mnt_idmap and from_vfsuid() is the identity. Link: https://patch.msgid.link/20260615-work-idmapped-overlayfs-v1-3-7381632aa402@kernel.org Reviewed-by: Amir Goldstein <amir73il@gmail.com> Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
2026-06-29ovl: handle idmapped mounts in ovl_permission()Christian Brauner1-1/+1
When the overlay mount is idmapped, the permission check on the overlay inode must account for the mount's idmapping. Use the struct mnt_idmap passed in by the VFS instead of the hardcoded &nop_mnt_idmap when checking the overlay inode against the caller's credentials. The second check, which verifies that the mounter may access the underlying real inode, continues to use the real layer's idmap mnt_idmap(realpath.mnt) under the mounter's credentials and is deliberately left unchanged. The overlay mount idmap only affects how the caller views the overlay inode, not the mounter's access to the layers, so it cannot widen access to the real files. No functional change until FS_ALLOW_IDMAP is set on ovl_fs_type; until then the overlay mount idmap is always &nop_mnt_idmap. Link: https://patch.msgid.link/20260615-work-idmapped-overlayfs-v1-2-7381632aa402@kernel.org Reviewed-by: Amir Goldstein <amir73il@gmail.com> Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
2026-06-29ovl: handle idmapped mounts in ovl_create_object() and ovl_tmpfile()Christian Brauner1-8/+8
In preparation for allowing the overlay mount itself to be idmapped, thread the mount's struct mnt_idmap into the inode creation path and use it for inode_init_owner() instead of the hardcoded &nop_mnt_idmap. The preallocated overlay inode's i_{u,g}id are copied into the override credentials by ovl_override_creator_creds() to create the real upper inode, so honoring the overlay mount idmap here makes newly created files, directories, special files, symlinks and tmpfiles get the caller's mapped fs{u,g}id once overlay mounts can be idmapped. No functional change: until FS_ALLOW_IDMAP is set on ovl_fs_type the overlay mount idmap is always &nop_mnt_idmap, so inode_init_owner() behaves exactly as before. Link: https://patch.msgid.link/20260615-work-idmapped-overlayfs-v1-1-7381632aa402@kernel.org Reviewed-by: Amir Goldstein <amir73il@gmail.com> Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
2026-06-29fs: look up the superblock via the device table in user_get_super()Christian Brauner1-20/+8
user_get_super() still finds the superblock for a device number by walking the global super_blocks list under sb_lock. Every superblock is registered in the device table under its s_dev since sget_fc() inserts it there, including superblocks on anonymous devices, so use the table instead. The refcount-pinning cursor helpers super_dev_{get,first,next}() only touch table state and do not depend on CONFIG_BLOCK, so drop the CONFIG_BLOCK guard around them: their new caller serves anonymous devices as well (ustat() on e.g. tmpfs) and is built without CONFIG_BLOCK. The guard falls in this patch rather than separately since without this caller the helpers would be unused without CONFIG_BLOCK. The pinned entry holds a passive reference on the superblock so super_lock() can be called directly; once the superblock is locked grab a passive reference for the caller before dropping the pin. The device table contains more than the old walk could find: a superblock is also registered for every additional device it claims (the xfs log and realtime devices, btrfs member devices, the ext4 external journal, erofs blob devices). Don't filter those out: specifying any device a filesystem uses now resolves to that filesystem, so ustat() and quotactl() work on e.g. the xfs log device or a btrfs member device (the latter used to fail outright as btrfs superblocks carry an anonymous s_dev that never matches a member device). When several superblocks share a device (erofs blob devices) the first live superblock wins. The cursor also keeps scanning past dying superblocks where the old walk gave up after the first s_dev match, so a mount racing with the unmount of the same device (or with the reuse of a recycled anonymous dev_t) finds the live superblock where the old walk could spuriously return NULL. This removes the last s_dev-keyed walk of the super_blocks list and takes ustat() and quotactl()'s block device lookup off sb_lock entirely. Link: https://patch.msgid.link/20260616-work-super-bdev_holder_global-v2-17-7df6b864028e@kernel.org Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
2026-06-29super: make fs_holder_ops privateChristian Brauner1-2/+1
Now that filesystems open and claim their block devices through fs_bdev_file_open_by_{dev,path}(), nothing outside fs/super.c references fs_holder_ops. Make it static and drop its declaration from blkdev.h. Reviewed-by: Jan Kara <jack@suse.cz> Link: https://patch.msgid.link/20260616-work-super-bdev_holder_global-v2-16-7df6b864028e@kernel.org Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
2026-06-29f2fs: open via dedicated fs bdev helpersChristian Brauner1-3/+3
Route the extra device opens of a multi-device f2fs through fs_bdev_file_open_by_path() so each device is registered against the superblock, and convert the matching release in destroy_device_list() to fs_bdev_file_release(). The first device aliases the main bdev file opened by setup_bdev_super() and is already registered through it. f2fs opened its extra devices without holder ops, so a freeze, sync, or removal of one of them was never propagated to the superblock. Registering them wires those events up: every device now freezes, thaws, syncs, and shuts down the filesystem like the main device does. Link: https://patch.msgid.link/20260616-work-super-bdev_holder_global-v2-15-7df6b864028e@kernel.org Acked-by: Chao Yu <chao@kernel.org> Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
2026-06-29erofs: open via dedicated fs bdev helpersChristian Brauner1-12/+23
Route opens through fs_bdev_file_open_by_path() so each external device is registered against the correct superblock, and convert the matching releases. Gao Xiang: I think typical immutable filesystems don't need .shutdown() and .remove_bdev() for the following reasons: - blk_mark_disk_dead() sets GD_DEAD in advance of fs_bdev_mark_dead() so that the following bios will fail immediately; block_device references are still valid so it seems overkill to handle dead blockdevs in the deep filesystem I/O submission path. - Immutable filesystems like EROFS don't have write paths and journals, so they don't need to block writes (i.e., new dirty pages), metadata changes, and abort journals. - The comment above loop_change_fd() documents a valid read-only use case we need to support anyway, but it calls disk_force_media_change() which will call fs_bdev_mark_dead() later: we don't want loop_change_fd() shutdowns the active filesystems and return -EIO unconditionally. Currently I think the default behavior (shrink_dcache_sb + evict_inodes) in fs_bdev_mark_dead() is enough for immutable filesystems, tried to document in the commit here for later reference. Signed-off-by: Gao Xiang <hsiangkao@linux.alibaba.com> Link: https://patch.msgid.link/20260616-work-super-bdev_holder_global-v2-14-7df6b864028e@kernel.org Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
2026-06-29fs: tolerate per-superblock freeze errors on shared devicesChristian Brauner1-3/+20
When several superblocks share a device, keep it frozen even if some of them failed to freeze and swallow the error: rolling the others back via thaw_super() can fail too, so neither is a clear win. A single filesystem still reports its error, and a sync_blockdev() failure is always reported. Thaw follows the same rule. A device can only be shared once superblocks claim it with a common exclusivity token, which erofs starts doing in the next patch; for everyone else the loop visits exactly one superblock and the behavior is unchanged. Link: https://patch.msgid.link/20260616-work-super-bdev_holder_global-v2-13-7df6b864028e@kernel.org Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
2026-06-29fs: look up superblocks via the device table in fs_holder_opsChristian Brauner2-138/+137
Switch the fs_holder_ops callbacks from recovering the single owning superblock out of bdev->bd_holder to walking the device-to-superblock table and acting on every superblock registered for the device. The holder argument becomes purely the block layer's exclusivity token and is no longer needed by the fs specific callbacks. All devices opened with fs_holder_ops are registered by now: the main device since setup_bdev_super() switched to fs_bdev_file_open_by_dev() and the extra devices (xfs log and realtime devices, btrfs member devices, the ext4 external journal) since the preceding per-filesystem conversions. So no event is lost in the switchover. The walk uses a refcount-pinning cursor: each step takes a reference on the entry via sd_ref and resumes from its sd_node. Unlinking an entry is deferred to the last unpin, so a cursor never resumes from a removed node. mark_dead and sync only need the passive reference the entry holds plus s_umount, which they take with super_lock_shared(). freeze and thaw additionally need an active reference and acquire it with get_active_super(), which waits for the superblock to be born before taking s_active. Taking s_active before the superblock is born would pin a still-mounting superblock so a racing mount that aborts could never drop s_active to zero and reach SB_DYING, deadlocking the wait for SB_BORN. This is how filesystems_freeze() and filesystems_thaw() acquire it too. One semantic change: when no live superblock uses the device anymore (the holder is dying or was never registered), fs_bdev_freeze() and fs_bdev_thaw() now return 0 - freeze after syncing the block device - where they used to return -EINVAL. The freeze-deny release path moves to the table in the same switchover. A device made unfreezable for a btrfs membership change must drop its table entry before re-allowing freezing; otherwise a freeze racing the release reaches the superblock through the still-registered entry and is stranded once the release unlinks it. Split fs_bdev_unregister() out of fs_bdev_file_release() - the inverse of fs_bdev_register() - so btrfs_release_device_allow_freeze() can drop the {dev, sb} entry, re-allow freezing on the still-open device, then close it. Re-allowing only after the entry is gone keeps a racing freeze from reaching the superblock, and doing it while the file is still open avoids touching the block device after the close. btrfs previously yielded bd_holder before re-allowing, which this commit makes irrelevant to freeze resolution. Link: https://patch.msgid.link/20260616-work-super-bdev_holder_global-v2-12-7df6b864028e@kernel.org Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
2026-06-29ext4: open via dedicated fs bdev helpersChristian Brauner1-6/+6
Route the external journal device open through fs_bdev_file_open_by_dev() so it is registered against the superblock, and convert the matching releases to fs_bdev_file_release(). Reviewed-by: Jan Kara <jack@suse.cz> Link: https://patch.msgid.link/20260616-work-super-bdev_holder_global-v2-11-7df6b864028e@kernel.org Tested-by: syzbot@syzkaller.appspotmail.com Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
2026-06-29btrfs: open via dedicated fs bdev helpersChristian Brauner1-9/+20
Route the device opens through fs_bdev_file_open_by_path() so each device is registered against the superblock, and convert the matching releases to fs_bdev_file_release(). The temporary identification opens that only read the superblock and close again pass a NULL holder and keep using bdev_fput(). On the close path the superblock is taken from bdev_file->private_data (the holder set at open) rather than device->fs_info->sb: a mount that fails before btrfs_init_devices_late() runs leaves device->fs_info NULL, which close_fs_devices() would otherwise dereference. Tested-by: syzbot@syzkaller.appspotmail.com Link: https://patch.msgid.link/20260616-work-super-bdev_holder_global-v2-10-7df6b864028e@kernel.org Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
2026-06-29xfs: port to fs_bdev_file_open_by_path()Christian Brauner2-6/+6
Route the log and rt device opens through fs_bdev_file_open_by_path() so each external device is registered against mp->m_super, and convert the matching releases to fs_bdev_file_release(). The data device is still opened and released by setup_bdev_super()/kill_block_super(); when the log lives on the data device the open resolves to the existing (dev, sb) entry so the superblock is acted on once. Reviewed-by: Jan Kara <jack@suse.cz> Link: https://patch.msgid.link/20260616-work-super-bdev_holder_global-v2-9-7df6b864028e@kernel.org Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
2026-06-29fs: add dedicated block device open helpers for filesystemsChristian Brauner3-10/+148
Add fs_bdev_file_open_by_{dev,path}() and fs_bdev_file_release(). They open the device with fs_holder_ops and register a claim in the device-to-superblock table. Claims on the same (device, superblock) pair share one entry, so when a filesystem claims a device it already uses (xfs with its log on the data device), no second entry is added and each superblock will be acted on once. The holder argument remains purely the block layer's exclusivity token: a superblock, or a file_system_type for a device shared by several superblocks of that type. The shared case only becomes usable once the fs_holder_ops callbacks resolve superblocks through the table instead of bdev->bd_holder. Convert the main device, setup_bdev_super() and kill_block_super(), over: the open finds the entry registered by sget_fc() and claims it again. cramfs and romfs bypass kill_block_super() so they can handle MTD mounts and release the main device with a plain bdev_fput(), which would leave the claim behind: the (dev, sb) entry would never be unregistered and the passive reference it holds would keep the superblock alive forever. Convert their release paths in the same step. The frozen-device check stays in setup_bdev_super() for the primary device and is added to fs_bdev_register() for new claims, i.e. every additional device a filesystem opens through the helpers. Only a (device, superblock) pair the superblock claimed earlier may be reopened while frozen (xfs with its log on the data device): the freeze already covers that superblock through the existing claim, so nothing escapes it. Without the setup_bdev_super() check a device frozen before the mount even started (dm lock_fs, loop) could be mounted and written to (journal replay) under an active freeze, because the primary open reuses the entry registered by sget_fc() and never takes the new-claim path. Both checks read bd_fsfreeze_count only after the entry is published (by sget_fc() for the primary, by fs_bdev_register() for new claims) and pair with bdev_freeze() incrementing the count before walking the table: either the mount sees the elevated freeze count and fails with EBUSY, or the freeze finds the published entry and converges once SB_BORN is set. Link: https://patch.msgid.link/20260616-work-super-bdev_holder_global-v2-8-7df6b864028e@kernel.org Reviewed-by: Jan Kara <jack@suse.cz> Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
2026-06-29fs: maintain a global device-to-superblock tableChristian Brauner3-2/+103
fs_holder_ops recovers the owning superblock from bdev->bd_holder, which forces the holder to be exactly one superblock and prevents several superblocks from sharing one block device. That's what erofs is doing. As a first step introduce a global dev_t-keyed rhltable mapping each device to the superblock(s) using it. The entry is preallocated in alloc_super() and registered under sb->s_dev by the set callback through set_anon_super() and set_bdev_super(), the two helpers every set callback assigns s_dev through. Registration is the final fallible act of a set callback, so an insert failure unwinds through sget_fc()'s existing set-failure path: the fs_context keeps ownership of s_fs_info and the callers' error paths stay correct. set_anon_super() releases the anonymous dev it allocated when registration fails. Unwinding through deactivate_locked_super() instead would run kill_sb() and free s_fs_info behind the caller's back: nfs and ceph free that object through a local pointer when sget_fc() fails and would double-free. The superblock stashes the entry in sb->s_super_dev and kill_super_notify() drops the claim through it, so teardown doesn't depend on s_dev staying stable; an entry that was never registered is freed together with the superblock in destroy_super_work(). Each table entry holds a passive reference (s_passive) on its superblock, so the struct stays valid for as long as the entry is reachable. Entries are claim-counted through sd_ref: additional claims on the same (device, superblock) pair share the entry, and the unlink is deferred to the last put, so a later iteration cursor never resumes from a removed node. The table is initialized from mnt_init(): the first superblocks (the tmpfs shm mount and rootfs) are created from start_kernel() long before any initcall runs, so an initcall would be too late. The table has no readers yet; the fs_holder_ops callbacks are switched over once all devices a filesystem claims are registered. Link: https://patch.msgid.link/20260616-work-super-bdev_holder_global-v2-7-7df6b864028e@kernel.org Reviewed-by: Jan Kara <jack@suse.cz> Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
2026-06-29ocfs2: don't reset s_dev on dismountChristian Brauner1-1/+0
ocfs2_dismount_volume() has reset sb->s_dev to zero since the original merge in ccd979bdbce9 ("[PATCH] OCFS2: The Second Oracle Cluster Filesystem") as part of scrubbing the super_block. Nothing reads the field afterwards: all ocfs2-internal uses are mount-time log and trace prints, and dev_t-keyed superblock lookups skip a dying superblock anyway - s_root is gone before ->put_super runs and super_lock() refuses SB_DYING superblocks. The upcoming device-to-superblock table registers every superblock under its s_dev. Drop the reset instead of leaving a superblock around whose s_dev contradicts its registration. Link: https://patch.msgid.link/20260616-work-super-bdev_holder_global-v2-6-7df6b864028e@kernel.org Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
2026-06-29ext4: use anonymous devices for KUnit test superblocksChristian Brauner2-14/+4
The mballoc and extents KUnit tests create superblocks through sget_fc() with a set callback that never assigns s_dev and a kill_sb that only calls generic_shutdown_super(). The upcoming global device-to-superblock table registers every superblock under its s_dev, so each superblock needs a unique device number. Allocate a proper anonymous device via set_anon_super_fc() and release it through kill_anon_super(). Link: https://patch.msgid.link/20260616-work-super-bdev_holder_global-v2-5-7df6b864028e@kernel.org Reviewed-by: Jan Kara <jack@suse.cz> Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
2026-06-29super: take lock after last reference countChristian Brauner1-35/+28
__put_super() required the caller to hold sb_lock, so put_super() wrapped it. The per-device superblock table introduced later drops its passive references from contexts that do not hold sb_lock, so make put_super() self-locking: drop the count first and take sb_lock only for the final list_del. With the count now dropped outside sb_lock a superblock can briefly sit on @super_blocks with s_passive == 0 before it is unlinked, so the list walkers (__iterate_supers(), iterate_supers_type(), user_get_super()) switch to refcount_inc_not_zero() and skip it. Link: https://patch.msgid.link/20260616-work-super-bdev_holder_global-v2-3-7df6b864028e@kernel.org Reviewed-by: Jan Kara <jack@suse.cz> Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
2026-06-29super: convert s_count to refcount_t s_passiveChristian Brauner1-9/+9
The superblock carries two counters: s_active, the active reference count that keeps the filesystem usable, and s_count, the passive reference count that merely keeps the structure itself alive. Turn the passive count into a refcount_t and rename it to s_passive to make the pairing with s_active obvious. Everything is still serialized by sb_lock, so there is no functional change; the conversion buys the usual refcount_t saturation and underflow checking. The following patches start dropping passive references without holding sb_lock and make the device-to-superblock table hold one passive reference per registered entry, which a plain integer cannot support. Link: https://patch.msgid.link/20260616-work-super-bdev_holder_global-v2-2-7df6b864028e@kernel.org Reviewed-by: Jan Kara <jack@suse.cz> Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
2026-06-29btrfs: deny freezing devices undergoing a replaceChristian Brauner3-14/+72
A device replace opens a target and, on success, frees the source on a live filesystem from btrfs_dev_replace_finishing() - which cannot fail and also runs from a kthread on mount resume. A bdev_freeze() racing the source free or the target swap-in would freeze the filesystem through a claim that is being torn down or replaced, leaving nothing for bdev_thaw() to rebalance. Make both devices unfreezable for the whole replace, with the invariant that a STARTED replace holds one deny on each device and any other state holds none. The target is denied at open (btrfs_open_device_deny_freeze(), undone on btrfs_init_dev_replace_tgtdev()'s error unwind); the source is denied at the start of btrfs_dev_replace_start(), before mark_block_group_to_copy() so every 'leave' unwind sees both denied. The deny tracks the STARTED state and is dropped whenever the replace leaves it: btrfs_dev_replace_finishing() re-allows the target it makes a member and frees the source through btrfs_close_bdev(allow_freeze=true), and its scrub-error path re-allows both as it cancels. Its early failures (before the device swap) keep the replace STARTED and resumable, so both stay denied. Suspending for unmount re-allows both, so they are reopened freezable at the next mount where btrfs_resume_dev_replace_async() re-denies them (staying suspended if a device is frozen right then); a replace cancelled from the suspended state therefore destroys the target without allowing. btrfs_close_bdev() and btrfs_destroy_dev_replace_tgtdev() take an allow_freeze argument to carry this distinction; the unmount path (btrfs_close_one_device()) passes false. On resume, a failed kthread_run() re-allows both devices and goes through the suspend path, resetting the replace to SUSPENDED and finishing the exclusive operation instead of returning straight away. The (re)mount still aborts on that error; routing it through suspend keeps the deny balanced against the unmount teardown and additionally drops BTRFS_EXCLOP_DEV_REPLACE, closing a pre-existing leak that was harmless on the failed mount that frees the fs but would have wedged future exclusive operations after a failed remount-rw. Link: https://patch.msgid.link/20260616-work-super-freeze_deny_upstream-v2-5-b3567c7f994b@kernel.org Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
2026-06-29btrfs: deny freezing a device while it is being addedChristian Brauner2-5/+43
btrfs_init_new_device() opens and claims the new device on a live superblock without holding the write count, so a bdev_freeze() racing the window between the claim being published and the device becoming a member could freeze the filesystem through a claim the add may still abort and tear down. Add btrfs_open_device_deny_freeze(): it opens the device once non-exclusively to take the freeze deny, then claims it by the same dev_t, so the holder is only ever published while the device is already unfreezable. Keep it denied until the add is durable: bdev_allow_freeze() on each success return (the device is now a committed member), btrfs_release_device_allow_freeze() on the error unwind. The deny spans the whole add, including the seeding tail whose late failures still release the device. A device already frozen when the add starts is refused with -EBUSY. Link: https://patch.msgid.link/20260616-work-super-freeze_deny_upstream-v2-4-b3567c7f994b@kernel.org Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
2026-06-29btrfs: deny freezing a device while it is being removedChristian Brauner3-2/+23
btrfs_rm_device() runs under mnt_want_write_file(), but the claim on the removed device is released by the ioctl after mnt_drop_write_file(), so a bdev_freeze() racing that window could freeze the filesystem through the device just as its claim is torn down, leaving nothing for bdev_thaw() to rebalance. The window cannot be closed by reordering the teardown. btrfs_rm_device() hands the final bdev_fput() back to the ioctl, run only after mnt_drop_write_file(), because bdev_release() takes the disk ->open_mutex and its dependency chain, which must not nest under the superblock's freeze/write protection -- freeze_super() drops s_umount before draining writers precisely to keep sb_start_write ordered above s_umount. Holding mnt_want_write across bdev_fput() would reintroduce that inversion, so the holder teardown is forced outside the write-protected section. A freeze landing in the resulting gap resolves the still-live holder, rides in, and strands when the claim is released; no ordering of the close against the drop removes the gap. The device itself therefore has to refuse freezing for the whole removal. Deny freezing the device for the duration of the removal: bdev_deny_freeze() at the start of btrfs_rm_device() (it cannot be frozen yet, the ioctl holds the write count), and release it through btrfs_release_device_allow_freeze() in the ioctls on success, or bdev_allow_freeze() on the error paths that keep the device a member. A device frozen before the removal begins is refused with -EBUSY. btrfs_release_device_allow_freeze() yields the holder, re-allows freezing, then closes the device, so the re-allow neither strands the filesystem on a racing freeze nor touches the block device after the final fput. Link: https://patch.msgid.link/20260616-work-super-freeze_deny_upstream-v2-3-b3567c7f994b@kernel.org Reviewed-by: Johannes Thumshirn <johannes.thumshirn@wdc.com> Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
2026-06-29fs: Add bpf_sock_read_xattr() kfunc to read socket xattrsChristian Brauner1-0/+37
In c8db08110cbe ("Merge tag 'vfs-7.1-rc1.xattr' of git://git.kernel.org/pub/scm/linux/kernel/git/vfs/vfs") we added support for extended attributes for sockets. This comes in two flavors: sockfs and non-sockfs/filesystem sockets. Filesystem sockets are actual filesystem objects so reading xattrs must use dedicated fs helpers such as bpf_get_dentry_xattr() and bpf_get_file_xattr(). Those are inherently sleeping operations. Sockfs sockets on the other hand don't need to use sleeping operations as the underlying data structure is lockless. In addition, retrieval of sockfs extended attributes often happens from LSM hooks that only provide struct socket and it's completely nonsensical to grab a reference to a file, then force a sleeping operation to retrieve the xattr and drop the reference. We know that the sockfs file cannot go away while the LSM hook runs. This series adds a bpf_sock_read_xattr() kfunc that, given a struct socket, reads a user.* extended attribute from the socket's sockfs inode into a bpf_dynptr. Together with fsetxattr() from userspace this lets a process label a socket with a user.* xattr and have a BPF LSM program retrieve that label locklessly. The kfunc mirrors the existing bpf_cgroup_read_xattr(), including the restriction to the user.* namespace. systemd uses user.* xattrs on sockets to implement socket rate limiting and to tag sockets for other purposes [1] such as implementing a varlink registry. There is currently no efficient way for a BPF program to read those labels back. The new helper allows a listening socket marked with an extended attribute to be read back during bind/connect and then act on the connect()ing socket. Extended attributes make it possible to allow an unprivileged user manager such as systemd --user to mark sockets from userspace and then rediscover them or implement policies. The kfunc is registered KF_RCU and only for BPF LSM programs. A struct socket is only guaranteed to live in sockfs when an LSM socket hook hands it out, which is what keeps SOCK_INODE() valid. Sockets that embed struct socket outside sockfs (tun, tap) are only reachable from tracing programs and are excluded by the registration. (Btw, for consistency it would be nice to force allocation of struct socket from sockfs instead of simply embedding it in e.g., struct tun_file which makes the SOCKFS_I() pattern a hazard - at least outside of sockfs functions.) The read never sleeps and takes no lock. For sockfs the value lives in the inode's in-memory xattr store and simple_xattr_get() resolves it with an RCU-protected rhashtable lookup, taking neither the inode lock nor any xattr lock. The kfunc is therefore usable from both sleepable and non-sleepable LSM hooks. Link: https://github.com/systemd/systemd/pull/40559 [1] Link: https://patch.msgid.link/20260617-work-bpf-sock-xattr-v1-1-a1276f7c9da3@kernel.org Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
2026-06-29efs: Remove EFSMatthew Wilcox (Oracle)11-1170/+0
The kernel EFS code has been unmaintained for over twenty years. It was superseded on IRIX around thirty years ago. I haven't seen an EFS filesystem in the wild since 1999. Userspace tools to read EFS filesystems exist, such as https://github.com/jkbenaim/efsextract There's no benefit to keeping this filesystem in the kernel, and it only increases the maintenance burden for tree-wide changes. Signed-off-by: Matthew Wilcox (Oracle) <willy@infradead.org> Link: https://patch.msgid.link/20260618211822.3599089-1-willy@infradead.org Acked-by: Jori Koolstra <jkoolstra@xs4all.nl> Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
2026-06-29fs: Free any excess xarray nodes in clear_inode()Matthew Wilcox (Oracle)1-13/+10
For many years we've had a hard to hit leak of xarray nodes. Hugh documented it well in commit 786b31121a2c. Recently people and syzbot have found ways to force it to happen with madvise. Rather than fix the leaks where they happen, just call xa_destroy() which has the side-effect of cycling the i_pages lock. Cc: Rik van Riel <riel@surriel.com> Cc: Zi Yan <ziy@nvidia.com> Cc: Jinjiang Tu <tujinjiang@huawei.com> Cc: Dave Jones <davej@codemonkey.org.uk> Link: https://lore.kernel.org/all/20260121062243.1893129-1-tujinjiang@huawei.com/ Signed-off-by: Matthew Wilcox (Oracle) <willy@infradead.org> Link: https://patch.msgid.link/20260623192850.1595958-1-willy@infradead.org Reviewed-by: Rik van Riel <riel@surriel.com> Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
2026-06-29fs/super: skip non-memcg-aware nr_cached_objects in memcg slab shrinkUsama Arif1-2/+17
The super_block shrinker is registered with SHRINKER_MEMCG_AWARE because its dentry and inode LRUs are memcg-aware (via list_lru). But the optional ->nr_cached_objects() hooks that the shrinker also drives are not memcg-aware: btrfs extent maps and xfs inode reclaim operate on filesystem-global state, and shmem's unused-huge shrinker walks a per-superblock shrinklist. None of them filter by sc->memcg. The mismatch shows up under memcg-heavy slab reclaim. shrink_slab_memcg() calls do_shrink_slab() once per (memcg, NUMA node) pair for every memcg whose bit is set in the per-superblock shrinker bitmap, which on a busy host means hundreds of calls per reclaim pass. Each scan queues the same global shrinker work item that's already kicked from the root path. Because btrfs/xfs global count is typically non-zero on any in-use filesystem, the returned total stays positive even if a memcg's own dentry/inode LRUs are empty. shrink_slab_memcg() therefore never clears the SB shrinker bit in the memcg bitmap, so subsequent reclaim passes from the same memcg re-enter super_cache_count() and pay for the global counter walk again. Restrict ->nr_cached_objects() to the global shrink path (sc->memcg NULL or root). The memcg-aware dentry/inode LRUs keep being counted and scanned per memcg as before; only the global fs-specific hooks are skipped. The root/global shrink path still drives those hooks; only their invocation from non-root memcg slab reclaim is removed. Signed-off-by: Usama Arif <usama.arif@linux.dev> Link: https://patch.msgid.link/20260609123047.1948242-1-usama.arif@linux.dev Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
2026-06-29fs: nullfs should include mount.hBen Dooks1-0/+2
The nullfs_fs_type is declared in mount.h but when declared in nullfs.c there is a warning as mount.h is not being included. Add include of "mount.h" to remove the following sparse warning: fs/nullfs.c:66:25: warning: symbol 'nullfs_fs_type' was not declared. Should it be static? Signed-off-by: Ben Dooks <ben.dooks@codethink.co.uk> Link: https://patch.msgid.link/20260617104721.900914-1-ben.dooks@codethink.co.uk Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
2026-06-29iomap: always return status from iomap_write_iterBrian Foster1-3/+1
iomap_write_iter() returns either an error code or 0 if partial progress has been made. The error sanitization was required in the past because the return value was used by iomap_iter() to determine how much progress was made, and thus what to pass to ->iomap_end() and how much to advance the iter. Now that iter handlers advance the iter incrementally and progress is separate from return status, this is no longer needed. iomap_iter() infers partial progress directly from the iter state and similarly, iomap_file_buffered_write() uses iter.pos to determine whether to return a short write or an error code. This also eliminates a minor quirk in the write iteration where if an error interrupts a partial write, we'd have to loop back into iomap_write_iter() once more and run into the error a second time before iomap terminates the operation and returns the error. With the error code returned directly and separate from write progress, we can complete the operation and return from iomap_iter() immediately. Signed-off-by: Brian Foster <bfoster@redhat.com> Link: https://patch.msgid.link/20260617114253.635751-1-bfoster@redhat.com Reviewed-by: Christoph Hellwig <hch@lst.de> Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
2026-06-29xfs: fix AGFL extent count calculation in xrep_agfl_filljiazhenyuan1-1/+1
In xrep_agfl_fill(), the call to xagb_bitmap_set() passes 'agbno - 1' as the length argument. However, xagb_bitmap_set() expects a length (number of blocks), not an end block number. Passing 'agbno - 1' causes used_extents to record an incorrect range. Fix this by calculating the correct length as 'agbno - start', which represents the actual number of blocks filled into the AGFL. Signed-off-by: jiazhenyuan <jiazhenyuan@uniontech.com> Fixes: 014ad53732d2ba ("xfs: use per-AG bitmaps to reap unused AG metadata blocks during repair") Reviewed-by: "Darrick J. Wong" <djwong@kernel.org> Signed-off-by: Carlos Maiolino <cem@kernel.org>
2026-06-29xfs: simplify the failure path in xfs_buf_alloc_vmallocChristoph Hellwig1-4/+3
Look at the __GFP_NORETRY flag set for readahead so that we don't have to pass both the gfp_t and the flags in. Signed-off-by: Christoph Hellwig <hch@lst.de> Reviewed-by: Carlos Maiolino <cmaiolino@redhat.com> Signed-off-by: Carlos Maiolino <cem@kernel.org>
2026-06-29xfs: fix incorrect use of gfp flags in xfs_buf_alloc_backing_memChristoph Hellwig1-26/+23
xfs_buf_alloc_backing_mem currently has two issues with how the GFP_ flags are set: - when aiming for a large folio allocation, the gfp mask is adjusted to try less hard, but these flags then persist for the vmalloc allocation, which is bogus. - the __GFP_NOFAIL for small allocations is also applied when readahead force __GFP_NORETRY which doesn't make any sense. Fix this by only applying __GFP_NOFAIL when __GFP_NORETRY is not set, and by reordering the code so that the large folio gfp adjustments are performed locally just for that allocation. Fixes: 94c78cfa3bd1 ("xfs: convert buffer cache to use high order folios") Signed-off-by: Christoph Hellwig <hch@lst.de> Reviewed-by: Carlos Maiolino <cmaiolino@redhat.com> Signed-off-by: Carlos Maiolino <cem@kernel.org>
2026-06-29xfs: lift setting __GFP_NOFAIL from xfs_buf_alloc_kmem to the callerChristoph Hellwig1-2/+2
The current __GFP_NOFAIL setting is wrong in some cases. Prepare for fixing that by giving control to the caller. Signed-off-by: Christoph Hellwig <hch@lst.de> Reviewed-by: Carlos Maiolino <cmaiolino@redhat.com> Signed-off-by: Carlos Maiolino <cem@kernel.org>
2026-06-29xfs: split up xfs_buf_alloc_backing_memChristoph Hellwig1-21/+40
Split out helpers for folio and vmalloc allocations to prepare for a bug fix. Signed-off-by: Christoph Hellwig <hch@lst.de> Reviewed-by: Carlos Maiolino <cmaiolino@redhat.com> Signed-off-by: Carlos Maiolino <cem@kernel.org>
2026-06-28debugfs: Fix lockdown check for mmap_prepareChun-Yi Lee1-1/+2
Commit 651fdda8406d ("relay: update relay to use mmap_prepare") changed the `mmap` file operation to `mmap_prepare` for relayfs, but the lockdown check in debugfs was not updated accordingly. This prevents debugfs from being locked down when the kernel is in integrity mode if a file uses `mmap_prepare` but not `mmap`. Since the conversion to `mmap_prepare` across the kernel is not yet complete, update the lockdown check to look for both `mmap` and `mmap_prepare` to ensure comprehensive coverage. Fixes: 651fdda8406d ("relay: update relay to use mmap_prepare") Signed-off-by: Chun-Yi Lee <jlee@suse.com> Cc: David Howells <dhowells@redhat.com> Cc: Lorenzo Stoakes <ljs@kernel.org> Cc: Andy Shevchenko <andy.shevchenko@gmail.com> Cc: Thomas Gleixner <tglx@linutronix.de> Cc: Rafael J. Wysocki <rafael@kernel.org> Cc: Matthew Garrett <mjg59@srcf.ucam.org> Cc: Greg Kroah-Hartman <gregkh@linuxfoundation.org> Cc: Danilo Krummrich <dakr@kernel.org> Cc: driver-core@lists.linux.dev Cc: linux-kernel@vger.kernel.org Cc: stable@vger.kernel.org Tested-by: Disha Goel <disgoel@linux.ibm.com> Reviewed-by: Lorenzo Stoakes <ljs@kernel.org> Link: https://patch.msgid.link/20260615104750.1000-1-jlee@suse.com Signed-off-by: Danilo Krummrich <dakr@kernel.org>
2026-06-27Merge tag 'fscrypt-for-linus' of git://git.kernel.org/pub/scm/fs/fscrypt/linuxLinus Torvalds4-216/+233
Pull fscrypt fixes from Eric Biggers: - Fix a bug where in a specific edge case, file contents en/decryption could be done with the wrong data unit size - Fix the data structure used for keeping track of users that have added an fscrypt key to be a simple list instead of a 'struct key' keyring This fixes issues such as a lockdep report found by syzbot and possible unintended interactions with the keyctl() system calls * tag 'fscrypt-for-linus' of git://git.kernel.org/pub/scm/fs/fscrypt/linux: fscrypt: Replace mk_users keyring with simple list fscrypt: Fix key setup in edge case with multiple data unit sizes
2026-06-26Merge tag 'ceph-for-7.2-rc1' of https://github.com/ceph/ceph-clientLinus Torvalds11-102/+1037
Pull ceph updates from Ilya Dryomov: "This adds support for manual client session reset in CephFS, allowing operators to get out of tricky livelock situations involving caps and file locks without evicting the problematic client instance on the MDS side or rebooting the client node both of which can be disruptive" * tag 'ceph-for-7.2-rc1' of https://github.com/ceph/ceph-client: ceph: add manual reset debugfs control and tracepoints ceph: add client reset state machine and session teardown ceph: add diagnostic timeout loop to wait_caps_flush() ceph: harden send_mds_reconnect and handle active-MDS peer reset ceph: use proper endian conversion for flock_len in reconnect ceph: convert inode flags to named bit positions and atomic bitops rbd: switch to dynamic root device
2026-06-26Merge tag 'gfs2-for-7.2' of ↵Linus Torvalds5-15/+57
git://git.kernel.org/pub/scm/linux/kernel/git/gfs2/linux-gfs2 Pull gfs2 updates from Andreas Gruenbacher: - fix page poisoning not handled correctly when growing files - quota initialization / destruction fixes: sleeping under a bitlock in PREEMPT_RT, broken quota_init error recovery, missing RCU synchronization * tag 'gfs2-for-7.2' of git://git.kernel.org/pub/scm/linux/kernel/git/gfs2/linux-gfs2: gfs2: page poisoning fix gfs2: Remove unused fallocate_chunk argument gfs2: fix use-after-free in gfs2_qd_dealloc gfs2: move quota_init qc iterator increment gfs2: fix quota init duplicate scan
2026-06-26Merge tag 'ecryptfs-7.2-rc1-updates' of ↵Linus Torvalds2-28/+12
git://git.kernel.org/pub/scm/linux/kernel/git/tyhicks/ecryptfs Pull ecryptfs updates from Tyler Hicks: "No functional changes, just code cleanups: - replace kmalloc()/snprintf() with kasprintf() - simplify code flow by removing an unnecessary variable" * tag 'ecryptfs-7.2-rc1-updates' of git://git.kernel.org/pub/scm/linux/kernel/git/tyhicks/ecryptfs: ecryptfs: use kasprintf in ecryptfs_crypto_api_algify_cipher_name ecryptfs: remove redundant variable found_auth_tok
2026-06-26Merge tag 'v7.2-rc-part2-smb3-server-fixes' of git://git.samba.org/ksmbdLinus Torvalds11-445/+1317
Pull smb server updates from Steve French: "This is mostly a correctness and compatibility update for ksmbd's SMB2/3 lease, oplock, durable handle, compound request, CREATE, rename, stream and share-mode handling. A large part of the series fixes cases found by smbtorture where ksmbd diverged from the SMB2/3 protocol requirements. The main changes are: - Rework SMB2 lease state handling so lease state is shared per ClientGuid/LeaseKey across opens, with better validation of lease create contexts, ACK handling, epochs, break-in-progress reporting, v2 lease notification routing, and chained lease breaks - Fix several oplock break corner cases, including ACK validation, timeout downgrade behavior, level-II break handling on unlink, share-conflict lease breaks, and read-control/stat-open behavior - Fix durable handle behavior around delete-on-close, stale reconnects, reconnect context parsing, oplock/lease break invalidation, and durable v2 AppInstanceId replacement - Fix compound request handling so related commands propagate failed statuses correctly, preserve response framing across chained errors, keep compound FIDs across READ/WRITE/FLUSH, and send interim STATUS_PENDING where clients expect cancellable compound I/O - Tighten CREATE and stream semantics, including create attribute validation, allocation size reporting, explicit create security descriptors, unnamed DATA stream handling, stream directory validation, and stream delete sharing against the base file - Fix rename and metadata behavior, including parent directory sharing checks, denying directory rename with open children, and preserving SMB ChangeTime across rename for open handles - Fix two important safety issues: a multichannel byte-range lock list owner race that could lead to use-after-free, and an NTLMv2 session key update before authentication proof validation - Fix a concurrent SMB2 NEGOTIATE preauth use-after-free, a UBSAN warning in compression capability parsing, a false hung-task warning in the durable handle scavenger, endian debug logging, Smatch indentation warnings, and kernel-doc warnings - Increase the default SMB3 transaction size from 1MB to 4MB to better match modern read/write negotiation and improve sequential I/O behavior" * tag 'v7.2-rc-part2-smb3-server-fixes' of git://git.samba.org/ksmbd: (50 commits) ksmbd: fix kernel-doc warnings in smb2_lease_break_noti() ksmbd: fix inconsistent indenting warnings ksmbd: validate NTLMv2 response before updating session key ksmbd: increase SMB3_DEFAULT_TRANS_SIZE from 1MB to 4MB ksmbd: fix UBSAN array-index-out-of-bounds in decode_compress_ctxt() ksmbd: sleep interruptibly in the durable handle scavenger ksmbd: start file id allocation at 1 ksmbd: treat read-control opens as stat opens only for leases ksmbd: validate :: stream type against directory create ksmbd: break conflicting-open leases only as far as needed ksmbd: break handle caching for share conflicts ksmbd: normalize ungrantable lease states ksmbd: return oplock protocol error for level II ack ksmbd: avoid level II oplock break notification on unlink ksmbd: downgrade oplock after break timeout ksmbd: apply create security descriptor first ksmbd: return requested create allocation size ksmbd: tighten create file attribute validation ksmbd: reject empty-attribute synchronize-only create ksmbd: honor stream delete sharing for base file ...
2026-06-25Merge tag 'v7.2-rc-part2-smb3-client-fixes' of ↵Linus Torvalds9-41/+117
git://git.samba.org/sfrench/cifs-2.6 Pull smb client fixes from Steve French: - fix potential double frees - fix potential memory leak in receiving compound response - querydir improvement - fix chown with smb311 posix extensions - ACL setting fixes - minor debug improvement and cleanup - add some missing protocol defines - sparse file fixes * tag 'v7.2-rc-part2-smb3-client-fixes' of git://git.samba.org/sfrench/cifs-2.6: cifs: define variable sized buffer for querydir responses smb/client: do not account EOF extension as allocation smb/client: preserve errors from smb2_set_sparse() smb: client: Fix next buffer leak in receive_encrypted_standard() smb/client: use %pe to print error pointer smb/client: name the default fallocate mode smb common: add missing AAPL defines smb/client: fix chown/chgrp with SMB3 POSIX Extensions smb/client: fix security flag calculation when setting security descriptors smb: client: refactor ACL setting control flow in id_mode_to_cifs_acl() smb: client: fix query directory replay double-free smb: client: fix change notify replay double-free smb: client: fix query_info() replay double-free smb: client: fix double-free in SMB2_close() replay smb: client: fix double-free in SMB2_ioctl() replay smb: client: fix double-free in SMB2_open() replay smb: client: fix double-free in SMB2_flush() replay
2026-06-25Merge tag 'net-7.2-rc1' of ↵Linus Torvalds2-2/+11
git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net Pull networking fixes from Jakub Kicinski: "Including fixes from netfilter and IPsec. Current release - regressions: - do not acquire dev->tx_global_lock in netdev_watchdog_up() - ethtool: keep rtnl_lock for ops using ethtool_op_get_link() - fix deadlock in nested UP notifier events Current release - new code bugs: - eth: - cn20k: fix subbank free list indexing for search order - airoha: fix BQL underflow in shared QDMA TX ring Previous releases - regressions: - netfilter: - flowtable: fix offloaded ct timeout never being extended - nf_conncount: prevent connlimit drops for early confirmed ct Previous releases - always broken: - require CAP_NET_ADMIN in the originating netns when modifying cross-netns devices - report NAPI thread PID in the caller's pid namespace - mac802154: fix dirty frag in in-place crypto for IOT radios - sctp: hold socket lock when dumping endpoints in sctp_diag, avoid an overflow - eth: gve: fix header buffer corruption with header-split and HW-GRO - af_key: initialize alg_key_len for IPComp states, prevent OOB read" * tag 'net-7.2-rc1' of git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net: (213 commits) selftests: bonding: add a test for VLAN propagation over a bonded real device vlan: defer real device state propagation to netdev_work net: add the driver-facing netdev_work scheduling API net: turn the rx_mode work into a generic netdev_work facility net: ethtool: keep rtnl_lock for ops using ethtool_op_get_link() rxrpc: Fix rxrpc_rotate_tx_rotate() to check there's something to rotate rxrpc: Fix leak of released call in recvmsg(MSG_PEEK) rxrpc: Fix socket notification race rxrpc: Fix potential infinite loop in rxrpc_recvmsg() rxrpc: Fix oob challenge leak in cleanup after notification failure rxrpc: Fix the reception of a reply packet before data transmission afs: Fix uncancelled rxrpc OOB message handler afs: Fix further netns teardown to cancel the preallocation charger rxrpc: Fix double unlock in rxrpc_recvmsg() rxrpc: Fix leak of connection from OOB challenge rxrpc: Fix ACKALL packet handling net: hns3: differentiate autoneg default values between copper and fiber net: hns3: fix permanent link down deadlock after reset net: hns3: refactor MAC autoneg and speed configuration net: hns3: unify copper port ksettings configuration path ...
2026-06-25afs: Fix uncancelled rxrpc OOB message handlerDavid Howells2-2/+6
Fix AFS to cancel its OOB message processing (typically to respond to security challenges). Also move OOB message processing to afs_wq so that it's also waited for and make the OOB handler just return if the net namespace is no longer live. Fixes: 5800b1cf3fd8 ("rxrpc: Allow CHALLENGEs to the passed to the app for a RESPONSE") Link: https://sashiko.dev/#/patchset/20260609140911.838677-1-dhowells%40redhat.com Signed-off-by: David Howells <dhowells@redhat.com> cc: Li Daming <d4n.for.sec@gmail.com> cc: Ren Wei <n05ec@lzu.edu.cn> cc: Marc Dionne <marc.dionne@auristor.com> cc: Jeffrey Altman <jaltman@auristor.com> cc: Simon Horman <horms@kernel.org> cc: linux-afs@lists.infradead.org cc: stable@kernel.org Link: https://patch.msgid.link/20260624163819.3017002-6-dhowells@redhat.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-06-25afs: Fix further netns teardown to cancel the preallocation chargerDavid Howells1-0/+5
When an afs network namespace is torn down, it cancels and waits for the work item that keeps the preallocated rxrpc call/conn/peer queue charged before disabling incoming (i.e. listen 0), but there's a small window in which it can be requeued by an incoming call wending through the I/O thread. Fix this by cancelling the charger work item again after reducing the listen backlog to zero. Fixes: 47694fbc9d24 ("afs: Fix netns teardown to cancel the preallocation charger") Reported-by: Jakub Kicinski <kuba@kernel.org> Signed-off-by: David Howells <dhowells@redhat.com> Link: https://sashiko.dev/#/patchset/20260609140911.838677-1-dhowells%40redhat.com cc: Li Daming <d4n.for.sec@gmail.com> cc: Ren Wei <n05ec@lzu.edu.cn> cc: Marc Dionne <marc.dionne@auristor.com> cc: Jeffrey Altman <jaltman@auristor.com> cc: Simon Horman <horms@kernel.org> cc: linux-afs@lists.infradead.org cc: stable@kernel.org Link: https://patch.msgid.link/20260624163819.3017002-5-dhowells@redhat.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-06-25ntfs: avoid stale runlist element dereference in MFT writebackCen Zhang1-2/+6
ntfs_write_mft_block() maps each $MFT record through the $MFT data runlist. For sub-folio clusters it looks up a struct runlist_element under ni->runlist.lock, drops the lock, and later uses rl->length and rl->vcn when choosing folio_sz. That pointer is only borrowed from ni->runlist.rl. Concurrent $MFT allocation extension can merge a replacement runlist under the same lock, and ntfs_rl_realloc() can free the old backing array. If that happens between the lookup and the later folio_sz decision, writeback can dereference freed runlist storage. The buggy scenario involves two paths, with each column showing the order within that path: MFT writeback path: $MFT allocation extension: 1. Look up rl under 1. Extend the $MFT data allocation. ni->runlist.lock. 2. Publish a replacement runlist. 2. Drop ni->runlist.lock. 3. Free the old runlist array. 3. Read rl->length and rl->vcn to choose folio_sz. Compute the remaining run length while ni->runlist.lock is still held, and use that scalar after unlock. This preserves the existing folio sizing decision without carrying a borrowed runlist_element across the lock boundary. Validation reproduced this kernel report: BUG: KASAN: slab-use-after-free in ntfs_mft_writepages+0x1c8d/0x1fb0 Call Trace: <TASK> dump_stack_lvl+0x66/0xa0 print_report+0xce/0x630 ? ntfs_mft_writepages+0x1c8d/0x1fb0 ? srso_alias_return_thunk+0x5/0xfbef5 ? __virt_addr_valid+0x20d/0x410 ? ntfs_mft_writepages+0x1c8d/0x1fb0 kasan_report+0xe0/0x110 ? ntfs_mft_writepages+0x1c8d/0x1fb0 ntfs_mft_writepages+0x1c8d/0x1fb0 ? __pfx_ntfs_mft_writepages+0x10/0x10 ? __pfx___mutex_unlock_slowpath+0x10/0x10 ? srso_alias_return_thunk+0x5/0xfbef5 ? iput+0x92/0xa80 do_writepages+0x219/0x530 ? __pfx_do_writepages+0x10/0x10 __writeback_single_inode+0x117/0xf50 ? do_raw_spin_lock+0x130/0x270 ? __pfx_do_raw_spin_lock+0x10/0x10 ? __pfx___writeback_single_inode+0x10/0x10 ? srso_alias_return_thunk+0x5/0xfbef5 writeback_sb_inodes+0x65b/0x1810 ? srso_alias_return_thunk+0x5/0xfbef5 ? lock_acquire+0x2b8/0x2f0 ? __pfx_writeback_sb_inodes+0x10/0x10 ? lock_release+0x1e0/0x280 ? _raw_spin_unlock+0x23/0x40 ? move_expired_inodes+0x2b8/0x850 __writeback_inodes_wb+0xf4/0x270 ? __pfx___writeback_inodes_wb+0x10/0x10 ? srso_alias_return_thunk+0x5/0xfbef5 ? queue_io+0x2e4/0x410 wb_writeback+0x666/0x880 ? srso_alias_return_thunk+0x5/0xfbef5 ? __pfx_wb_writeback+0x10/0x10 ? srso_alias_return_thunk+0x5/0xfbef5 ? srso_alias_return_thunk+0x5/0xfbef5 ? get_nr_dirty_inodes+0x1c/0x170 wb_workfn+0x75e/0xbb0 ? srso_alias_return_thunk+0x5/0xfbef5 ? _raw_spin_unlock_irqrestore+0x27/0x60 ? __pfx_wb_workfn+0x10/0x10 ? __pfx_debug_object_deactivate+0x10/0x10 ? lock_acquire+0x2b8/0x2f0 ? srso_alias_return_thunk+0x5/0xfbef5 ? lock_release+0x1e0/0x280 process_one_work+0x8d0/0x1870 ? __pfx_process_one_work+0x10/0x10 ? srso_alias_return_thunk+0x5/0xfbef5 worker_thread+0x575/0xf80 ? __pfx_worker_thread+0x10/0x10 kthread+0x2e7/0x3c0 ? __pfx_kthread+0x10/0x10 ret_from_fork+0x576/0x810 ? __pfx_ret_from_fork+0x10/0x10 ? srso_alias_return_thunk+0x5/0xfbef5 ? __switch_to+0x57e/0xe10 ? __switch_to_asm+0x33/0x70 ? __pfx_kthread+0x10/0x10 ret_from_fork_asm+0x1a/0x30 </TASK> Allocated by task 970: kasan_save_stack+0x33/0x60 kasan_save_track+0x14/0x30 __kasan_kmalloc+0xaa/0xb0 __kvmalloc_node_noprof+0x353/0x920 ntfs_rl_realloc+0x3c/0x80 ntfs_runlists_merge+0x1212/0x3010 ntfs_mft_data_extend_allocation_nolock+0x3e0/0x1f40 ntfs_mft_record_alloc+0x1ab4/0x4f10 __ntfs_create+0x680/0x2e50 ntfs_create+0x1e6/0x3a0 path_openat+0x2b55/0x3c10 do_file_open+0x1f4/0x460 do_sys_openat2+0xde/0x170 __x64_sys_openat+0x122/0x1e0 do_syscall_64+0x115/0x6a0 entry_SYSCALL_64_after_hwframe+0x77/0x7f Freed by task 1294: kasan_save_stack+0x33/0x60 kasan_save_track+0x14/0x30 kasan_save_free_info+0x3b/0x60 __kasan_slab_free+0x5f/0x80 kfree+0x307/0x580 ntfs_rl_realloc+0x66/0x80 ntfs_runlists_merge+0x1212/0x3010 ntfs_mft_data_extend_allocation_nolock+0x3e0/0x1f40 ntfs_mft_record_alloc+0x1ab4/0x4f10 __ntfs_create+0x680/0x2e50 ntfs_create+0x1e6/0x3a0 path_openat+0x2b55/0x3c10 do_file_open+0x1f4/0x460 do_sys_openat2+0xde/0x170 __x64_sys_openat+0x122/0x1e0 do_syscall_64+0x115/0x6a0 entry_SYSCALL_64_after_hwframe+0x77/0x7f Fixes: 115380f9a2f9 ("ntfs: update mft operations") Assisted-by: Codex:gpt-5.5 Signed-off-by: Cen Zhang <zzzccc427@gmail.com> Reviewed-by: Hyunchul Lee <hyc.lee@gmail.com> Signed-off-by: Namjae Jeon <linkinjeon@kernel.org>
2026-06-24cifs: define variable sized buffer for querydir responsesShyam Prasad N4-3/+17
QueryDirectory responses today are stored in one of two fixed sized buffers: smallbuf (448 bytes) or bigbuf (16KB). These are borrowed from server struct and are not sufficient for large-sized query dir operations. With this change we will now define a new buffer type specifically for cifs_search_info to hold variable sized responses. These will be allocated by kmalloc and freed by kfree. Signed-off-by: Shyam Prasad N <sprasad@microsoft.com> Signed-off-by: Steve French <stfrench@microsoft.com>
2026-06-24smb/client: do not account EOF extension as allocationHuiwen He1-3/+10
cifs_setsize() updates the local inode size after SetEOF succeeds. It also used the new EOF as a local i_blocks estimate, but extending EOF does not prove that the intervening range was allocated. For example, after writing 1 MiB and then extending EOF to 10 MiB, the client can report the file as fully allocated even though the server still reports a much smaller AllocationSize: $ dd if=/dev/zero of=test bs=1M count=1 $ truncate -s 10M test && stat -c 'size=%s blocks=%b' test $ stat --cached=never -c 'size=%s blocks=%b' test client stat: size=10485760 blocks=20480 server stat: size=10485760 blocks=2056 client stat after revalidation: size=10485760 blocks=2056 A later attribute revalidation may correct i_blocks, but callers such as xfstests generic/495 invoke swapon immediately after truncate. The swapfile hole check can therefore observe the inflated local i_blocks value and accept a sparse file. Do not grow i_blocks from cifs_setsize() on EOF extension. Only clamp it on shrink; allocation growth must come from write completion or from server-reported AllocationSize. With this change, EOF extension no longer makes a sparse file appear fully allocated before the next attribute revalidation, and xfstests generic/495 no longer accepts it through the inflated local i_blocks value. Signed-off-by: Huiwen He <hehuiwen@kylinos.cn> Reviewed-by: ChenXiaoSong <chenxiaosong@kylinos.cn> Signed-off-by: Steve French <stfrench@microsoft.com>
2026-06-24Merge tag 'ntfs3_for_7.2' of ↵Linus Torvalds13-302/+590
https://github.com/Paragon-Software-Group/linux-ntfs3 Pull ntfs3 updates from Konstantin Komarov: "Added: - depth limit to indx_find_buffer() to prevent stack overflow - validate split-point offset in indx_insert_into_buffer() - bounds check to run_get_highest_vcn() - fileattr_get() and fileattr_set() support - zero stale pagecache beyond valid data length - handle delayed allocation overlap in run lookup - validate lcns_follow in log_replay() conversion - cap RESTART_TABLE free-chain walker at rt->used - resize log->one_page_buf when adopting on-disk page size - reject direct userspace writes to reserved $LX* xattrs Fixed: - out-of-bounds read in decompress_lznt() - avoid -Wmaybe-uninitialized warnings - hold ni_lock across readdir metadata walk - preserve non-DOS attribute bits in system.dos_attrib - validate index entry key bounds - syncing wrong inode on DIRSYNC cross-directory rename - validate Dirty Page Table capacity in log_replay() copy_lcns - wrong LCN in run_remove_range() when splitting a run - allocate iomap inline_data using alloc_page - mount failure on 64K page-size kernels - out-of-bounds read in ntfs_dir_emit() and hdr_find_e() - bound attr_off in UpdateResidentValue against data_off - bound DeleteIndexEntryAllocation memmove length - bound copy_lcns dp->page_lcns[] index in analysis pass - bound NTFS_DE view.data_off in UpdateRecordData{Root,Allocation} - prevent potential lcn remains uninitialized Changed: - bound to_move in indx_insert_into_root() before hdr_insert_head() - call _ntfs_bad_inode() when failing to rename - fold resident writeback into writepages loop - force waiting for direct I/O completion - fold file size handling into ntfs_set_size() - reject SEEK_DATA and SEEK_HOLE past EOF early - format code, add descriptive comments and remove non-useful" * tag 'ntfs3_for_7.2' of https://github.com/Paragon-Software-Group/linux-ntfs3: (34 commits) ntfs3: reject direct userspace writes to reserved $LX* xattrs fs/ntfs3: resize log->one_page_buf when adopting on-disk page size fs/ntfs3: prevent potential lcn remains uninitialized ntfs3: cap RESTART_TABLE free-chain walker at rt->used fs/ntfs3: bound NTFS_DE view.data_off in UpdateRecordData{Root,Allocation} fs/ntfs3: validate lcns_follow in log_replay conversion fs/ntfs3: bound copy_lcns dp->page_lcns[] index in analysis pass fs/ntfs3: bound DeleteIndexEntryAllocation memmove length fs/ntfs3: bound attr_off in UpdateResidentValue against data_off ntfs3: fix out-of-bounds read in ntfs_dir_emit() and hdr_find_e() fs/ntfs3: fix mount failure on 64K page-size kernels ntfs3: avoid another -Wmaybe-uninitialized warning ntfs3: Allocate iomap inline_data using alloc_page fs/ntfs3: format code, deal with comments fs/ntfs3: reject SEEK_DATA and SEEK_HOLE past EOF early fs/ntfs3: fold file size handling into ntfs_set_size() fs/ntfs3: force waiting for direct I/O completion fs/ntfs3: fold resident writeback into writepages loop fs/ntfs3: handle delayed allocation overlap in run lookup fs/ntfs3: zero stale pagecache beyond valid data length ...
2026-06-24smb/client: preserve errors from smb2_set_sparse()Huiwen He1-15/+15
smb2_set_sparse() converts every FSCTL_SET_SPARSE failure to false and marks sparse support as broken for the share. Callers therefore report EOPNOTSUPP even for errors such as ENOSPC or EACCES, and later sparse operations remain disabled until the share is unmounted. Return the SMB2 ioctl error directly. Set broken_sparse_sup only for EOPNOTSUPP, which indicates that the server does not support the request. Update smb3_punch_hole() to propagate the error returned by smb2_set_sparse(). Fixes: 3d1a3745d8ca ("Add sparse file support to SMB2/SMB3 mounts") Signed-off-by: Huiwen He <hehuiwen@kylinos.cn> Reviewed-by: ChenXiaoSong <chenxiaosong@kylinos.cn> Signed-off-by: Steve French <stfrench@microsoft.com>
2026-06-24smb: client: Fix next buffer leak in receive_encrypted_standard()Haoxiang Li1-4/+6
receive_encrypted_standard() allocates next_buffer before checking whether the number of compound PDUs already reached MAX_COMPOUND. If the limit check fails, the function returns immediately and the newly allocated next_buffer is not assigned to server->smallbuf/server->bigbuf, making it leaked. Move the MAX_COMPOUND check before allocating next_buffer. Fixes: b24df3e30cbf ("cifs: update receive_encrypted_standard to handle compounded responses") Cc: stable@vger.kernel.org Signed-off-by: Haoxiang Li <haoxiang_li2024@163.com> Signed-off-by: Steve French <stfrench@microsoft.com>