summaryrefslogtreecommitdiff
AgeCommit message (Collapse)AuthorFilesLines
2026-08-07ACPI: APEI: Handle repeated SEA error stormsJunhao He1-1/+9
When hardware memory corruption occurs and a user process accesses the corrupted page, the CPU triggers a Synchronous External Abort (SEA). The kernel invokes do_sea() to handle the exception, which calls memory_failure() to handle the faulty page. Scenario 1: Memory Error Interrupt First, then SEA The page is already poisoned by the memory error interrupt path. The subsequent SEA handler sends a SIGBUS to the task, which accesses the poisoned page. This flow is correct. Scenario 2: SEA first, then memory error interrupt (problematic scenario) If a user task directly accesses corrupted memory through a PFNMAP-style mapping (e.g., devmem), the page may still be in the free-buddy state when SEA is handled. In this case, memory_failure() will poison the page without invoking kill_accessing_process(), and then takes the free-buddy recovery path. After the CPU returns to the task context, the task re-enters the SEA handler due to the same access. However, ghes_estatus_cached() suppresses all subsequent entries during the 10-second window, preventing ghes_do_proc() from being called. This suppression blocks the MF_ACTION_REQUIRED-based SIGBUS delivery, causing the kernel to fail to kill the task immediately. Consequently, the process keeps re-entering the SEA handler, leading to an SEA storm. Later, the memory error interrupt path also cannot kill the task, leaving the system stuck in this repeated loop. The following error logs are explained using the devmem process: NOTICE: SEA Handle [Hardware Error]: Hardware error from APEI Generic Hardware Error Source: 9 [Hardware Error]: event severity: recoverable [Hardware Error]: section_type: ARM processor error [Hardware Error]: physical fault address: 0x0000001000093c00 [T54990] Memory failure: 0x1000093: recovery action for free buddy page: Recovered [ T9955] EDAC MC0: 1 UE Multi-bit ECC on unknown memory (page:0x1000093 offset:0xc00 grain:1 - APEI location: ...) NOTICE: SEA Handle NOTICE: SEA Handle ... ... ---> SEA storm ... NOTICE: SEA Handle [ T9955] Memory failure: 0x1000093: already hardware poisoned ghes_print_estatus: 1 callbacks suppressed [Hardware Error]: Hardware error from APEI Generic Hardware Error Source: 9 [Hardware Error]: event severity: recoverable [Hardware Error]: section_type: ARM processor error [Hardware Error]: physical fault address: 0x0000001000093c00 [T54990] Memory failure: 0x1000093: already hardware poisoned [T54990] 0x1000093: Sending SIGBUS to devmem:54990 due to hardware memory corruption To resolve this, return an error when encountering the same SEA again. The subsequent SEA handler invocation uses arm64_notify_die() to send a SIGBUS signal to the task, which terminates the process and prevents it from re-entering the handler loop. Signed-off-by: Junhao He <hejunhao3@h-partners.com> Reviewed-by: Wupeng Ma <mawupeng1@huawei.com> Reviewed-by: Shuai Xue <xueshuai@linux.alibaba.com> Reviewed-by: Tony Luck <tony.luck@intel.com> Link: https://patch.msgid.link/20260527082707.2013499-1-hejunhao3@h-partners.com Signed-off-by: Rafael J. Wysocki <rafael.j.wysocki@intel.com>
2026-08-07ACPI: APEI: Fix ERST timeout unit conversionNirmoy Das1-1/+1
The ACPI specification defines bits 63:32 returned by GET_EXECUTE_OPERATION_TIMINGS as the maximum execution time in microseconds. erst_get_timeout() instead multiplies the value by NSEC_PER_MSEC. Use NSEC_PER_USEC to express the firmware-provided microsecond timeout in the nanosecond units expected by erst_timedout(). Fixes: fac475aab70b ("ACPI: APEI: Use ERST timeout for slow devices") Cc: stable@vger.kernel.org Signed-off-by: Nirmoy Das <nirmoyd@nvidia.com> Reviewed-by: Hanjun Guo <guohanjun@huawei.com> Link: https://patch.msgid.link/20260721182551.2434933-1-nirmoyd@nvidia.com Signed-off-by: Rafael J. Wysocki <rafael.j.wysocki@intel.com>
2026-08-07fwctl/bnxt: Add DMA buffer support for HWRM commandsPavan Chebbi2-5/+395
Several HWRM commands carry __le64 DMA address fields in their input structures; firmware reads from or writes to the memory those addresses point to. Have a static per-command descriptor table in the driver that records the details of the DMA fields in each supported HWRM input struct. When a DMA-bearing HWRM command arrives, the driver reads the userspace pointer out of each declared address field and clears the field, allocates a DMA-coherent kernel buffer sized from the command's own length information, copies data to/from the userspace pointer, and patches the field with the real DMA bus address before the command is sent to firmware. Responses are copied back to the original userspace pointer afterward. Scope-gated allow-list and timeout value list are updated with the new the commands. Signed-off-by: Pavan Chebbi <pavan.chebbi@broadcom.com> Link: https://patch.msgid.link/20260807125846.45570-3-pavan.chebbi@broadcom.com Signed-off-by: Jason Gunthorpe <jgg@nvidia.com>
2026-08-07bnxt_en: Update bnxt firmware specPavan Chebbi1-0/+585
Since bnxt_fwctl is going to support additional commands in the next patch, add their missing definitions from the firmware spec. Signed-off-by: Pavan Chebbi <pavan.chebbi@broadcom.com> Link: https://patch.msgid.link/20260807125846.45570-2-pavan.chebbi@broadcom.com Signed-off-by: Jason Gunthorpe <jgg@nvidia.com>
2026-08-07btrfs: skip hole detection during full fsync for files without holesFilipe Manana1-0/+9
If we the no-holes feature is enabled (a default since btrfs-progs 5.15), when doing a full fsync we always iterate of all leaves in the subvolume root that contain file extent items in order to detect holes between them. This can take a lot of time for files with a large number of extents. But if we know there are no prealloc extents and the amount of space (uncompressed space) is greater than or equals to the i_size of the inode, then we cannot have holes and therefore avoid searching for them. So skip the search if those conditions are met. The following test script was used: $ cat test.sh #!/bin/bash MNT=/mnt/nullb0 DEV=/dev/nullb0 umount $MNT &> /dev/null mkfs.btrfs -f $DEV mount $DEV $MNT # 256M gives 64K extents of 4K each. FILE_SIZE=$((256 * 1024 * 1024)) touch $MNT/foobar for ((i = 0; i < $FILE_SIZE; i += 8192)); do xfs_io -c "pwrite -S 0xab $i 4K" $MNT/foobar > /dev/null done xfs_io -c "fsync" $MNT/foobar for ((i = 4096; i < $FILE_SIZE; i += 8192)); do xfs_io -c "pwrite -S 0xab $i 4K" $MNT/foobar > /dev/null done # unmount and mount, clear caches and ensure the next fsync is a # full sync. umount $MNT mount $DEV $MNT # Do some change to the file in order to fsync. xfs_io -c "pwrite -S 0xcd 0 4K" $MNT/foobar > /dev/null T0=$(date +%s%N) xfs_io -c "fsync" $MNT/foobar T1=$(date +%s%N) echo echo "Took $(( (T1 - T0) / 1000 ))us" umount $MNT Before this change: Took 28721us After this change: Took 5453us That's about 5.3x times faster. Reviewed-by: Qu Wenruo <wqu@suse.com> Signed-off-by: Filipe Manana <fdmanana@suse.com> Signed-off-by: David Sterba <dsterba@suse.com>
2026-08-07btrfs: add extra ASSERT()s to make sure the folio size is correctQu Wenruo1-0/+19
Inspired by the previous crash exposed by generic/795, we want to make sure every folio from btrfs page cache is properly aligned to block size. This is especially important for bs > ps support, as every btrfs infrastructure, e.g. extent map and extent state, requires strong block alignment checks. Furthermore, also output the minimal folio order from the inode mapping, which is the determining factor during debugging, helping a lot pinning down the final cause. Reviewed-by: Filipe Manana <fdmanana@suse.com> Signed-off-by: Qu Wenruo <wqu@suse.com> Signed-off-by: David Sterba <dsterba@suse.com>
2026-08-07btrfs: use GFP_NOWAIT for tree block readaheadBoris Burkov1-1/+2
extent_buffer readahead should not be able to painfully stall a search_slot and hog tree locks by getting stuck in direct reclaim. If the allocation fails, that is fine, we simply fail to do the readahead in that case. Reviewed-by: Jeff Layton <jlayton@kernel.org> Reviewed-by: Filipe Manana <fdmanana@suse.com> Signed-off-by: Boris Burkov <boris@bur.io> Reviewed-by: David Sterba <dsterba@suse.com> Signed-off-by: David Sterba <dsterba@suse.com>
2026-08-07btrfs: enable unlocked NOFAIL retry for eb allocationsBoris Burkov5-14/+50
Now that we have the btrfs_eb_prealloc struct to carry the allocation and the "needs prealloc" signal, wire that up between the various search_slot style callers down into alloc_extent_buffer. If the prealloc struct indicates that it supports a nowait try, then alloc_extent_buffer tries to allocate NOWAIT. If that succeeds, great. Otherwise, we return EAGAIN and signal via the struct that preallocation is required. The caller then does the allocation and tries again with the eb, bfs, and folios wired through in the prealloc struct. If unlock-and-allocate retries are not supported then we just use the normal gfp flags like before. Note that there are still two GFP_NOFS allocations, as far as I know, that happen under the lock and cannot be preallocated: - the __xa_cmpxchg to insert the eb into the eb xarray - the xarray allocations for filemap_add_folio to add the folios to the btree_inode mapping. The former we could wire up with xa_reserve if we signaled the "prealloc start" back up to the retry point. However, since there is no concept of reservation in the filemap xarray, it seemed relatively unhelpful to bother. These allocations are relatively small cached slab allocations, so hopefully we can move the needle on reclaim stalls without reserving them. Reviewed-by: Jeff Layton <jlayton@kernel.org> Reviewed-by: Filipe Manana <fdmanana@suse.com> Signed-off-by: Boris Burkov <boris@bur.io> Reviewed-by: David Sterba <dsterba@suse.com> Signed-off-by: David Sterba <dsterba@suse.com>
2026-08-07btrfs: add struct btrfs_eb_preallocBoris Burkov7-67/+166
In further preparation for supporting NOFAIL allocations with retries outside the critical section, add a struct to carry the extent_buffer and btrfs_folio_state we need to allocate. Refactor the allocation pathways to use the new struct but with no functional change. Wire empty prealloc structs in from callers. Reviewed-by: Filipe Manana <fdmanana@suse.com> Reviewed-by: Jeff Layton <jlayton@kernel.org> Signed-off-by: Boris Burkov <boris@bur.io> Reviewed-by: David Sterba <dsterba@suse.com> Signed-off-by: David Sterba <dsterba@suse.com>
2026-08-07btrfs: factor init_extent_buffer from __alloc_extent_bufferBoris Burkov1-8/+12
In preparation for preallocating extent_buffer data, factor eb initialization away from specifically allocating it. This allows us to allocate the eb, bfs, folios, etc. together in the main search_slot code paths, but still share initialization code with the dummy/test/clone allocation paths. Reviewed-by: Filipe Manana <fdmanana@suse.com> Signed-off-by: Boris Burkov <boris@bur.io> Reviewed-by: David Sterba <dsterba@suse.com> Signed-off-by: David Sterba <dsterba@suse.com>
2026-08-07btrfs: qgroup: fix a wrong length calculation in qgroup_free_reserved_data()Qu Wenruo1-7/+11
In that function, we round down the start position and round up the ending position. But during the calculation of @len, we use "round_up(start + len, sectorsize)", which is the rounded up end position, not the rounded up length. Which results a much larger length, and later we are still using "start + len", which is completely incorrect. Fix it by declaring a local @aligned_start and @aligned_len and use them instead. Fixes: bc42bda22345 ("btrfs: qgroup: Fix qgroup reserved space underflow by only freeing reserved ranges") Reviewed-by: Filipe Manana <fdmanana@suse.com> Signed-off-by: Qu Wenruo <wqu@suse.com> Signed-off-by: David Sterba <dsterba@suse.com>
2026-08-07btrfs: add validation for extent statesQu Wenruo1-0/+21
Extent maps have the extra validation since commit 3f255ece2f1e ("btrfs: introduce extra sanity checks for extent maps"), but extent states do not have a similar check. Introduce a basic alignment check for the following call sites, so that we can cover all extent states inserted into the tree: - insert_state_fast() - insert_state() - split_state() Reviewed-by: Filipe Manana <fdmanana@suse.com> Signed-off-by: Qu Wenruo <wqu@suse.com> Signed-off-by: David Sterba <dsterba@suse.com>
2026-08-07btrfs: use aligned range for locking in reflinkQu Wenruo1-2/+2
In btrfs_extent_same_range() and btrfs_clone_files(), the range passed into btrfs_lock_extent() is not aligned at its end, because we can reflink until the EOF, which may not be block aligned. Although this is not a big deal, for the sake of consistency, and to prepare for the upcoming stricter alignment check, pass an aligned range end to btrfs_lock_extent() and btrfs_unlock_extent(). Reviewed-by: Filipe Manana <fdmanana@suse.com> Signed-off-by: Qu Wenruo <wqu@suse.com> Signed-off-by: David Sterba <dsterba@suse.com>
2026-08-07btrfs: use aligned range for locking in extent_fiemap()Qu Wenruo1-2/+2
The @end parameter for all extent io tree helpers is inclusive, but the call site in extent_fiemap() is passing an exclusive end into btrfs_lock_extent(), which will step into the next block unexpectedly. Pass the inclusive end into btrfs_lock_extent() and btrfs_unlock_extent(). Fixes: ac3c0d36a2a2 ("btrfs: make fiemap more efficient and accurate reporting extent sharedness") Reviewed-by: Filipe Manana <fdmanana@suse.com> Signed-off-by: Qu Wenruo <wqu@suse.com> Signed-off-by: David Sterba <dsterba@suse.com>
2026-08-07btrfs: zoned: don't clobber the extent buffer when zeroing it outJohannes Thumshirn2-13/+31
On a zoned filesystem a freed-but-still-dirty tree block is written out as zeros (EXTENT_BUFFER_ZONED_ZEROOUT) only to keep the zone write pointer advancing. btree_csum_one_bio() implemented this by memzeroing the extent buffer's own folios before submission. That destroys the in-memory buffer while it may still be referenced. In particular btrfs_free_tree_block() can run on it afterwards and reads the header to add a delayed reference; once the header has been zeroed it frees bytenr 0 and corrupts the extent tree (the btrfs_header_bytenr(buf) != 0 ASSERT in btrfs_free_tree_block(), or an "unable to find ref" abort). It is flaky and reproduces under fsstress, e.g. generic/461 and generic/013. Write the zeros to disk from the shared zero page instead and leave the extent buffer content untouched, so any later reference - including the delayed reference from btrfs_free_tree_block() - still sees a valid header. end_bbio_meta_write() now clears writeback on the buffer's own folios, as the bio no longer carries them. Fixes: aa6313e6ff2b ("btrfs: zoned: don't clear dirty flag of extent buffer") Assisted-by: LLM (debugging, commit message) Reviewed-by: Boris Burkov <boris@bur.io> Signed-off-by: Johannes Thumshirn <johannes.thumshirn@wdc.com> Signed-off-by: David Sterba <dsterba@suse.com>
2026-08-07btrfs: zoned: drop stranded dirty metadata buffers at unmountJohannes Thumshirn1-0/+7
On a zoned filesystem a freed tree block is kept dirty and flagged EXTENT_BUFFER_ZONED_ZEROOUT so a later writeback zeroes it out and advances the zone write pointer. Unsynced tree-log updates (e.g. from rename or link) leave such buffers behind when the log is freed at commit, and across log generations they can end up ahead of the write pointer behind a hole, so btree_writepages() can never write them. During normal operation the space is later reclaimed by a zone reset; at unmount it is not, and the buffers survive to the final iput() of the btree inode, which hangs in folio_wait_writeback() once the endio workqueues are stopped. They cannot be written back from where they are freed (free_log_tree(), inside the committing transaction) without deadlocking against that commit, and they are stale anyway, not referenced by the committed superblock. Drop their dirty state in close_ctree(), before btrfs_stop_all_workers(). Signed-off-by: Johannes Thumshirn <johannes.thumshirn@wdc.com> Signed-off-by: David Sterba <dsterba@suse.com>
2026-08-07btrfs: zoned: drop stranded dirty metadata on transaction abortJohannes Thumshirn3-16/+60
On a zoned filesystem a freed tree block is not cleared but kept dirty and flagged EXTENT_BUFFER_ZONED_ZEROOUT, so a later writeback zeroes it out and advances the zone write pointer. A transaction abort turns the filesystem read-only before that writeback runs, so these buffers stay dirty and stranded ahead of the write pointer where btree_writepages() can no longer write them. They survive to the final iput() of the btree inode at unmount, which submits the write after the endio workqueues are gone, hanging unmount in folio_wait_writeback(). Clear the dirty state of such buffers when cleaning up the aborted transaction, where the buffer tree still references all of them. Assisted-by: LLM (debugging, commit message) Reviewed-by: Boris Burkov <boris@bur.io> Signed-off-by: Johannes Thumshirn <johannes.thumshirn@wdc.com> Signed-off-by: David Sterba <dsterba@suse.com>
2026-08-07btrfs: zoned: flush active metadata block group at btree_writepages() startJohannes Thumshirn1-21/+92
btree_writepages() writes the btree inode's dirty metadata in ascending logical address order. On a zoned filesystem only one metadata and one system block group is active for writing at a time, and check_bg_is_active() (via btrfs_check_meta_write_pointer()) pivots the active block group as writeback moves from one block group to the next. If the active block group sits at a higher logical address than another block group that also holds dirty metadata, the ascending walk reaches the lower one first and, to write it, has to finish the active block group and activate the lower one. It cannot finish a block group that still has unsent IO, and during WB_SYNC_ALL && !for_sync (commit) writeback it deliberately refuses to wait for that IO under fs_info->zoned_meta_io_lock, as that can deadlock. The pivot thus cannot issue the submission itself either, so it gives up: btrfs_check_meta_write_pointer() returns -EAGAIN, which btrfs_write_and_wait_transaction() treats as fatal and aborts the transaction, forcing the filesystem read-only. This happens intermittently under metadata-heavy relocation (e.g. fstests btrfs/187). Flush the active metadata and system block groups at the start of btree_writepages(), under the fs_info->zoned_meta_io_lock it already holds, so they have no unsent IO left and the later pivot can finish them and make forward progress. Fixes: 13bb483d32ab ("btrfs: zoned: activate metadata block group on write time") Assisted-by: LLM (debugging, commit message) Reviewed-by: Boris Burkov <boris@bur.io> Signed-off-by: Johannes Thumshirn <johannes.thumshirn@wdc.com> Signed-off-by: David Sterba <dsterba@suse.com>
2026-08-07btrfs: convert reflink.c to use btrfs_inode as parametersQu Wenruo1-52/+49
Inside reflink.c we still have a lot of functions passing VFS inode pointers, then internally convert them into btrfs_inode pointers. For example, inside btrfs_clone(), we have 12 BTRFS_I() call sites, while only 3 callsites that really require a VFS inode pointer. Do the cleanup to convert the following functions to pass a btrfs_inode pointer instead of a vanilla inode pointer: - btrfs_clone() - btrfs_extent_same_range() - clone_finish_inode_update(). Which covers all ad-hoc BTRFS_I() call sites inside reflink.c. Reviewed-by: Daniel Vacek <neelx@suse.com> Signed-off-by: Qu Wenruo <wqu@suse.com> Reviewed-by: David Sterba <dsterba@suse.com> Signed-off-by: David Sterba <dsterba@suse.com>
2026-08-07btrfs: use simple booleans for log_commit field in struct btrfs_rootFilipe Manana5-19/+13
We are using atomic types for the log_commit array of struct btrfs_root but all we need is simple booleans. The log_commit array elements are always protected by the root's log_mutex, both for writes and reads, so we can use a simple boolean. The use of atomics if from the very early days of the log tree code where the access to the fields was not protected by any lock. So switch to simple booleans, which results in cheaper code and slightly reduces the object size too. Reviewed-by: Boris Burkov <boris@bur.io> Reviewed-by: Jeff Layton <jlayton@kernel.org> Signed-off-by: Filipe Manana <fdmanana@suse.com> Signed-off-by: David Sterba <dsterba@suse.com>
2026-08-07btrfs: check for exit condition after waking in wait_log_commit()Filipe Manana1-4/+4
We check for the exit condition after we add ourselves to the wait queue and before we unlock the root's log_mutex, sleep and lock again log_mutex. This is not incorrect, but it's not optimal since in the first iteration this is pointless because we already know that root->log_commit[index] is not zero, so we should check the exit condition only after unlocking log_mutex, sleeping, waking up and locking again the log_mutex. So move the check for the exit condition to bottom of the loop, after we were woken and locked log_mutex again. Reviewed-by: Boris Burkov <boris@bur.io> Reviewed-by: Jeff Layton <jlayton@kernel.org> Signed-off-by: Filipe Manana <fdmanana@suse.com> Signed-off-by: David Sterba <dsterba@suse.com>
2026-08-07btrfs: move condition for log commit wait into wait_log_commit()Filipe Manana1-26/+17
Instead of having every caller check for root->log_commit[] being non-zero and then call wait_log_commit(), move the check into wait_log_commit() and have the callers call it unconditionally. Reviewed-by: Boris Burkov <boris@bur.io> Reviewed-by: Jeff Layton <jlayton@kernel.org> Signed-off-by: Filipe Manana <fdmanana@suse.com> Signed-off-by: David Sterba <dsterba@suse.com>
2026-08-07btrfs: remove log batch counter use for fsyncFilipe Manana4-14/+1
We have the log batch counter defined per root which is now useless after the previous patch (titled: "btrfs: stop sleeping for one jiffy in non-ssd mounts during log commit"). The counter is incremented early in the fsync path, before and after flushing dellaloc and waiting for writeback, and then the counter is read during the log sync path. The goal was to wait for tasks that are about to join a log transaction, so that we could reduce the amount of IO and log syncing (flush all log tree extent buffers and write super blocks), but that mechanism does not work since if there are currently no log writers, btrfs_sync_log() does not unlock the root's log_mutex, so no new log writers can join the log transaction. Having concurrent fsync tasks increasing the log_batch counter only makes us loop unnecessarily in btrfs_sync_log() - that is always true since the previous patch mentioned above and was true before that patch only when not using the "-o ssd" mount option (which is activated by default if the filesystem does not have rotational devices). So remove the log batch counter. No performance changes were observed after removing it. Reviewed-by: Boris Burkov <boris@bur.io> Reviewed-by: Jeff Layton <jlayton@kernel.org> Signed-off-by: Filipe Manana <fdmanana@suse.com> Signed-off-by: David Sterba <dsterba@suse.com>
2026-08-07btrfs: stop sleeping for one jiffy in non-ssd mounts during log commitFilipe Manana2-19/+1
Joining/starting a log transaction tracks if we ever had more than one task concurrently logging by setting the flag BTRFS_ROOT_MULTI_LOG_TASKS in the respective root. Once set, this flag remains for the rest of the lifetime of the transaction, only cleared when we don't have a log root and need to create a new one (transaction commits drop log roots). During log commit, if we are not on a ssd mount (or use the -o nossd mount option) and the BTRFS_ROOT_MULTI_LOG_TASKS flag is set, we sleep for one jiffy with the excuse to allow future log writers to join and log inodes and then commit a larger log transaction to reduce overall IO. However this is extremely inefficient because: 1) If at some point we had multiple tasks logging concurrently but now we have only one task at a time, we force it to wait for 1 jiffy; 2) One jiffy can vary between 1ms to 10ms, depending on the kernel config option CONFIG_HZ, which by default has a value of 250HZ and that corresponds to 4ms - that is a lot. This massively reduces the latency of fsyncs for non-ssd mounts, even on consumer grade spinning disks. Remove this mechanism to track if we have (or ever had) multiple tasks logging and wait for 1 jiffy. The following fio test was used to benchmark: $ cat fio-buffered-fsync.sh DEV=/dev/sdj MNT=/mnt/sdj MOUNT_OPTIONS="" MKFS_OPTIONS="" if [ $# -ne 6 ]; then echo "Use $0 NUM_JOBS FILE_SIZE IO_SIZE FSYNC_FREQ BLOCK_SIZE [write|randwrite]" exit 1 fi NUM_JOBS=$1 FILE_SIZE=$2 IO_SIZE=$3 FSYNC_FREQ=$4 BLOCK_SIZE=$5 WRITE_MODE=$6 if [ "$WRITE_MODE" != "write" ] && [ "$WRITE_MODE" != "randwrite" ]; then echo "Invalid WRITE_MODE, must be 'write' or 'randwrite'" exit 1 fi cat <<EOF > /tmp/fio-job.ini [writers] rw=$WRITE_MODE fsync=$FSYNC_FREQ fallocate=none group_reporting=1 direct=0 bs=$BLOCK_SIZE ioengine=psync filesize=$FILE_SIZE io_size=$IO_SIZE directory=$MNT numjobs=$NUM_JOBS EOF echo echo "Using config:" echo cat /tmp/fio-job.ini echo umount $MNT &> /dev/null mkfs.btrfs -f $MKFS_OPTIONS $DEV mount $MOUNT_OPTIONS $DEV $MNT fio /tmp/fio-job.ini umount $MNT Running the script as: ./fio-buffered-fsync.sh 8 64M 64M 1 4K randwrite Before patch: WRITE: bw=2647KiB/s (2711kB/s), 2647KiB/s-2647KiB/s (2711kB/s-2711kB/s), io=512MiB (537MB), run=198055-198055msec After patch: WRITE: bw=14.9MiB/s (15.6MB/s), 14.9MiB/s-14.9MiB/s (15.6MB/s-15.6MB/s), io=512MiB (537MB), run=34471-34471msec That's about 5.7 times faster. Reviewed-by: Boris Burkov <boris@bur.io> Reviewed-by: Jeff Layton <jlayton@kernel.org> Signed-off-by: Filipe Manana <fdmanana@suse.com> Signed-off-by: David Sterba <dsterba@suse.com>
2026-08-07btrfs: sysfs: fix path of the "read_policy" module parameter in commentZenghui Yu1-1/+1
The correct path of the "read_policy" module parameter should be /sys/module/btrfs/parameters/read_policy. Fix it. Acked-by: Randy Dunlap <rdunlap@infradead.org> Signed-off-by: Zenghui Yu <zenghui.yu@linux.dev> Reviewed-by: David Sterba <dsterba@suse.com> Signed-off-by: David Sterba <dsterba@suse.com>
2026-08-07btrfs: retry verity reads for not-uptodate Merkle foliosYichong Chen1-5/+11
btrfs_read_merkle_tree_page() can find a folio in the mapping that is not uptodate. After taking the folio lock, the current code treats that state as a read error and returns -EIO. That can make a previous transient read failure sticky. If the failed read left a not-uptodate folio in the mapping, later callers find that folio and fail instead of retrying the read. Keep the existing page-cache insertion and locking order, but retry the Merkle item read when a not-uptodate folio is found in the mapping. Also unlock the folio when read_key_bytes() fails so that a later caller can lock it and retry the read. Fixes: 06ed09351b67 ("btrfs: convert btrfs_read_merkle_tree_page() to use a folio") Reviewed-by: Boris Burkov <boris@bur.io> Signed-off-by: Yichong Chen <chenyichong@uniontech.com> Signed-off-by: David Sterba <dsterba@suse.com>
2026-08-07btrfs: use %pe for error code outputQu Wenruo12-93/+94
During an interrupted mount, I got the following messages: workqueue: Failed to create a rescuer kthread for wq "btrfs-qgroup-rescan": -EINTR BTRFS error (device dm-3): open_ctree failed: -12 Workqueue code is outputting a human readable error string, meanwhile we're still using a numeric error code. So follow the workqueue code to use "%pe" format, which will automatically convert an error pointer to the human readable string. However this is a minor pitfall, if the return value is not an error code, e.g. a positive number, "%pe" with "ERR_PTR(ret)" will output the pointer as a hash value, e.g.: ret=1 %pe out=0000000019414716 ret=-22 %pe out=-EINVAL So we should not use this "%pe" output for callsites that are known to return positive values. Reviewed-by: Johannes Thumshirn <johannes.thumshirn@wdc.com> Reviewed-by: Jeff Layton <jlayton@kernel.org> Signed-off-by: Qu Wenruo <wqu@suse.com> Signed-off-by: David Sterba <dsterba@suse.com>
2026-08-07btrfs: open code BTRFS_BYTES_TO_BLKS()Qu Wenruo2-6/+4
That macro is only utilized 4 times, all inside file.c, while we have tons of open-coded usages. And since it's a macro, there is no proper type checks at all. There isn't much need for such a rarely utilized macro. Signed-off-by: Qu Wenruo <wqu@suse.com> Reviewed-by: David Sterba <dsterba@suse.com> Signed-off-by: David Sterba <dsterba@suse.com>
2026-08-07btrfs: replace writeback inhibition xarray with a fixed inline bufferLeo Martins3-37/+77
Commit f9a48549a15a ("btrfs: inhibit extent buffer writeback to prevent COW amplification") tracks the extent buffers a transaction handle has inhibited in a per-handle xarray. Keying the tracking to the transaction handle is correct, but using an xarray for it causes two problems in production. First, a write_iops regression. Every COW calls btrfs_inhibit_eb_writeback() from btrfs_force_cow_block() and should_cow_block(), which does an xa_store() keyed by eb->start. The kernel test robot reported a 22.6% fio.write_iops regression on a single-task 4k randwrite workload (ftruncate ioengine, buffered IO) on btrfs. The cost is the per-COW xarray store done on every COW'd block. Replacing it with a non-allocating fixed buffer recovers the lost throughput, and that buffer does more per-COW bookkeeping yet still recovers, so the cost is the xarray operation itself rather than the extra tracking work. Second, an unbounded cleanup walk. btrfs_uninhibit_all_eb_writeback() iterates every eb the handle inhibited with xa_for_each(). A single handle that COWs a very large number of blocks (inode eviction, or truncate of a file with many extents, where btrfs_truncate_inode_items() loops over many search_again descents under one handle) makes that walk arbitrarily long. It runs in __btrfs_end_transaction() before num_writers is dropped, so it blocks the committing thread; this shows up as multi-second stalls and RCU stall reports. Replace the xarray with a fixed inline array on btrfs_trans_handle, managed with a CLOCK (second-chance) eviction policy. Inhibiting a buffer becomes an array append with no allocation and no tree walk, and the end-of-handle cleanup is bounded by the array size. The set that actually needs protection is the working set the handle revisits across search_again descents, the search path frontier, which is on the order of the tree height. It is not every block the handle ever COWs. should_cow_block() re-inhibiting an already tracked buffer marks it referenced, so revisited buffers survive eviction while write-once buffers are reclaimed first. A small fixed buffer is therefore enough where a non-evicting array would either overflow or have to grow without bound. BTRFS_INHIBITED_EBS_SLOTS is 8 and the reference bits pack into a u32. The CLOCK eviction is what justifies the extra complexity over a plain non-evicting array. The test workload stresses amplification: it removes 16 heavily fragmented 64 MiB files in one transaction while background writeback keeps writing out in-use metadata. A re-COW event is a buffer already COWed in the running transaction that was written back and then COWed again; the figure below is the ratio of re-COW events to first-COW events summed across the eviction (n=5, lower is better): tracking re-COW per first-COW no inhibition 6.1 non-evicting array, 32 slots 3.8 CLOCK array, 8 slots (this patch) 1.6 unbounded xarray (reverted) 1.4 The non-evicting array fills with write-once buffers and stops covering the buffers the handle keeps revisiting, so even at four times the slots it leaves most of the amplification. CLOCK evicts the cold buffers and keeps the revisited ones, recovering almost all of the unbounded benefit. The eviction policy, not the buffer size, is what closes the gap. eb->writeback_inhibitors and the WB_SYNC_ALL bypass in lock_extent_buffer_for_io() are unchanged, so fsync and commit behavior are unaffected. A reference is taken on each tracked buffer so it cannot be freed while the array points at it; eviction drops that reference and the inhibitor count. There's another testing report, showing 20% latency improvement on reflink and deduplication synthetic benchmark. Full detailed report at https://github.com/lcf0399/linux-regression-evidence/tree/main/btrfs-remap-writeback-inhibition-v2 . Link: https://lore.kernel.org/all/CANGjgd=fQkHht2PdDi-+EAdzWH7UtxxWhhJ7b80Rr17PbpgxOw@mail.gmail.com/ Reported-by: kernel test robot <oliver.sang@intel.com> Fixes: f9a48549a15a ("btrfs: inhibit extent buffer writeback to prevent COW amplification") Closes: https://lore.kernel.org/oe-lkp/202603112240.f7605968-lkp@intel.com Tested-by: Chengfeng Lin <lin2530632123@gmail.com> Reviewed-by: Filipe Manana <fdmanana@suse.com> Reviewed-by: Sun YangKai <sunk67188@gmail.com> Signed-off-by: Leo Martins <loemra.dev@gmail.com> Reviewed-by: David Sterba <dsterba@suse.com> Signed-off-by: David Sterba <dsterba@suse.com>
2026-08-07btrfs: compression: allocate heuristic buckets with workspaceRosen Penev1-12/+2
Avoid allocating the heuristic buckets separately from the workspace, the lifetime is the same. The new size of struct heuristic_ws is 2112. SLUB merges same/similar sized structures for the named caches, so there's a chance such size already exists on the system, like below: $ grep 2112 /proc/slabinfo sighand_cache 593 1335 2112 15 8 Signed-off-by: Rosen Penev <rosenp@gmail.com> Reviewed-by: David Sterba <dsterba@suse.com> Signed-off-by: David Sterba <dsterba@suse.com>
2026-08-07btrfs: check if root is readonly when setting posix aclSun YangKai1-0/+4
For a filesystem which has btrfs read-only property set to true, all write operations including acl and xattr should be denied. However, acl can still be set even if btrfs ro property is true. This happens because no function on the set_acl code path checks the root is readonly or not. It was checked in btrfs_setxattr_trans() but got removed in commit 353c2ea735e4 ("btrfs: remove redundant readonly root check in btrfs_setxattr_trans") That commit didn't check if all the callers properly check the root's read-only flag. A previous fix is commit b51111271b03 ("btrfs: check if root is readonly while setting security xattr"). Always check if the root is read-only before performing the set acl operation. Fixes: 353c2ea735e4 ("btrfs: remove redundant readonly root check in btrfs_setxattr_trans") Reviewed-by: Johannes Thumshirn <johannes.thumshirn@wdc.com> Reviewed-by: Filipe Manana <fdmanana@suse.com> Signed-off-by: Sun YangKai <sunyangkai@fnnas.com> Signed-off-by: Filipe Manana <fdmanana@suse.com> Signed-off-by: David Sterba <dsterba@suse.com>
2026-08-07btrfs: drop recovered reloc root refs on recovery failureGuanghui Yang1-4/+27
During relocation recovery, each fs root gets a reference to its relocation root. If loading or adding a later root fails, or if the first transaction commit fails, btrfs_recover_relocation() jumps to out_unset before merge_reloc_roots() and clean_dirty_subvols(). put_reloc_control() drops the list-owned relocation root references, but it does not clear fs_root->reloc_root or drop the references owned by those pointers. Mount cleanup only drops them when BTRFS_FS_ERROR is set, so an error such as -ENOMEM while processing a later root can leave references behind. Keep temporary references to the fs roots associated during recovery. On failure, clear their reloc_root pointers and drop the corresponding references. Once the first transaction commit succeeds, drop only the temporary fs root references and let the normal merge and cleanup paths handle the relocation roots. Fault injection on a pending-relocation image confirmed the cleanup gap. With an injected first-commit failure, 25 fs roots had reloc_root set with fs_error=0. With this fix, the same failure path drops that count to 0 before mount fails. Fixes: f44deb7442ed ("btrfs: hold a ref on the root->reloc_root") CC: stable@vger.kernel.org Signed-off-by: Guanghui Yang <3497809730@qq.com> Reviewed-by: David Sterba <dsterba@suse.com> Signed-off-by: David Sterba <dsterba@suse.com>
2026-08-07btrfs: add missing sctx check in cleanup path in btrfs_ioctl_send()Hongling Zeng1-1/+1
Add sctx NULL check in the for loop condition of the sort_clone_roots cleanup path for consistency with the else branch. Reviewed-by: Boris Burkov <boris@bur.io> Signed-off-by: Hongling Zeng <zenghongling@kylinos.cn> Reviewed-by: David Sterba <dsterba@suse.com> Signed-off-by: David Sterba <dsterba@suse.com>
2026-08-07btrfs: disguise single-data-RAID56 as RAID1/RAID1C3Qu Wenruo3-4/+27
Recently kernel RAID56 lib is trying to remove the unexpected single-data-RAID56 (2 disks RAID5 or 3 disk RAID5) support, meanwhile btrfs still supports such setup, which means in the long run btrfs has to handle such corner case by ourselves. Thankfully single-data-RAID56 is really RAID1/RAID1C3, since data and P/Q stripes all match each other, rotation also makes no difference. This patch will disguise those single-data-RAID56 chunks as RAID1/RAID1C3 chunks. This is done at two timings: - Chunk read - Chunk allocation This is done by introducing btrfs_chunk_map::on_disk_type member, which stores the type read from the on-disk metadata. Meanwhile btrfs_chunk_map::type is calculated using on_disk_type. For most profiles @type matches @on_disk_type, but for single-data-RAID56, the @type will be RAID1/RAID1C3. This method has a minimal impact on the fs, all other operations like scrub and read-repair, are all based on the chunk map type, so the disguise method will require no extra modification to those call sites. Although there are still some locations that are checking against block_group->flags, e.g. scrub. Those call sites will still get extra limits assuming the bg is RAID56. But it should not cause any extra problem. Signed-off-by: Qu Wenruo <wqu@suse.com> Signed-off-by: David Sterba <dsterba@suse.com>
2026-08-07btrfs: remove duplicated block group type assignmentQu Wenruo1-1/+0
In the function fill_dummy_bgs(), bg->flags is assigned twice. Just remove the second assignment. Signed-off-by: Qu Wenruo <wqu@suse.com> Reviewed-by: David Sterba <dsterba@suse.com> Signed-off-by: David Sterba <dsterba@suse.com>
2026-08-07btrfs: remove btrfs_chunk_map::io_(align|width) membersQu Wenruo2-6/+0
Those two members are read from on-disk metadata, but never utilized. And for new chunks we always set those members to BTRFS_STRIPE_LEN anyway. Thus there is no need to keep them inside btrfs_chunk_map. Signed-off-by: Qu Wenruo <wqu@suse.com> Signed-off-by: David Sterba <dsterba@suse.com>
2026-08-07btrfs: make sure EXTENT_BUFFER_READING is cleared under refs_lockQu Wenruo2-12/+47
[FALSE ALERTS] There is a bug report that the warning inside invalidate_and_check_btree_folios() got triggered during btrfs/298: BTRFS info (device sdd): first mount of filesystem f9bf732a-a19b-44b9-99a7-614ddff168e2 BTRFS info (device sdd): using crc32c checksum algorithm BTRFS error (device sdd): failed to find fsid cb2fdb42-b638-4f2f-badd-4127467ba674 when attempting to open seed devices BTRFS error (device sdd): failed to read chunk tree: -2 ------------[ cut here ]------------ WARNING: disk-io.c:3342 at invalidate_and_check_btree_folios+0x260/0x3c0 [btrfs], CPU#4: mount/125993 CPU: 4 UID: 0 PID: 125993 Comm: mount Tainted: G W OE 7.1.0-rc7-custom+ #1 PREEMPT(full) Hardware name: QEMU KVM Virtual Machine, BIOS edk2-20250812-19.fc42 08/12/2025 Call trace: invalidate_and_check_btree_folios+0x260/0x3c0 [btrfs] (P) open_ctree+0x1f50/0x23b0 [btrfs] btrfs_get_tree+0x89c/0xc48 [btrfs] vfs_get_tree+0x30/0x110 vfs_cmd_create+0x58/0xe8 __arm64_sys_fsconfig+0x39c/0x518 invoke_syscall.constprop.0+0x48/0x120 el0_svc_common.constprop.0+0x40/0xe8 do_el0_svc+0x24/0x38 el0_svc+0x50/0x310 el0t_64_sync_handler+0xa0/0xe8 el0t_64_sync+0x198/0x1a0 ---[ end trace 0000000000000000 ]--- BTRFS warning (device sdd): unable to release extent buffer 365985792 owner 3 gen 17 refs 3 flags 0x5 [CAUSE] In that invalidate_and_check_btree_folios() we wait for the eb to finish its read, then check if it's only held by us and the btree inode. If not, then do a warning as it may be still held, and could cause problems. But there is a small window where the check can lead to false alerts: Thread A (Read endio) | Thread B (Unmount) ----------------------------------+------------------------------------- end_bbio_meta_read() | | The eb has one extra ref held | | by the reader, and has | | EXTENT_BUFFER_READING flag set | invalidate_and_check_btree_folios() | | | |- clear_extent_buffer_reading() | | | | |- wait_on_bit_io(); | | | The EXTENT_BUFFER_READING flag is | | | cleared | | |- if (refcount_read(eb->refs) > 2) | | The eb is held by the read, us | | and btree inode, thus it | | will trigger the warning |- free_extent_buffer() | [FIX] Introduce a helper, free_extent_buffer_clear_reading(). If the new parameter, @clear_reading, is set, we will hold the spinlock at the beginning of free_extent_buffer_clear_reading() to make sure the EXTENT_BUFFER_READING flag is cleared inside the same critical section of decreasing refs. Now free_extent_buffer() will just call free_extent_buffer_clear_reading() with @clear_reading set to false, so no behavior change. But for end_bbio_meta_read(), it will not clear_extent_buffer_reading() directly, but pass @clear_reading as true. Then inside invalidate_and_check_btree_folios(), hold the refs_lock before reading refs. So that we eliminate the race window completely. Reported-by: Su Yue <glass.su@suse.com> Link: https://lore.kernel.org/linux-btrfs/DC0C775E-13B3-47D9-9AB2-895BB11C029D@suse.com/ Fixes: 83f7e52b7ed1 ("btrfs: warn about extent buffer that can not be released") Reviewed-by: Filipe Manana <fdmanana@suse.com> Signed-off-by: Qu Wenruo <wqu@suse.com> Signed-off-by: David Sterba <dsterba@suse.com>
2026-08-07btrfs: write-protect folios during data writebackBoris Burkov3-6/+53
commit 095be159f3eb ("btrfs: unify folio dirty flag clearing") replaced the folio_clear_dirty_for_io() call in extent_write_cache_pages() with a plain folio_test_dirty() check. Besides clearing the dirty flag, folio_clear_dirty_for_io() also calls folio_mkclean(), which write-protects the shared mmap PTEs mapping the folio. Note that we still do call folio_clear_dirty_for_io() later in submit_one_sector() when we clear dirty on the last sector of the folio (the only sector for non-subpage cases). But we lost this early call in extent_write_cache_pages(). Without the extra write-protection, a process with the file mmap-ed can modify a sector while it is being used by writeback in a way that expects a stable folio (checksumming, compressing, copying, etc...) without faulting, which manifests as a handful of concrete bugs. 1. For large folios or subpage sectorsize, it is possible to submit a bio which does not cover the whole folio. When this happens, we will have a bio in flight for a folio that we have *not* called folio_clear_dirty_for_io() on. If a task with an existing mmap-ed PTE writes (without faulting..) in this window, it can result in corruptions. If the write arrives while the checksumming or writing itself is underway, this can result in an invalid checksum and later corruption reports on read. If the write arrives after checksumming/writing is done but before the last sector dirty is cleared, then the write is present in page cache but doesn't affect the dirty tracking and will be lost when the folio is fully finished being submitted and the dirty bit is cleared. This results in losing the write even if fsync() is called. 2. For zoned submissions which are done in batch separate from the main extent_writepage() loop, we also risk csum violations for those submissions. Zoned writes are clamped to max_zone_append_size and are not aligned with folios, so a submission can span two folios. The first folio being processed in extent_write_cache_pages() will call extent_write_locked_range() which will submit the partial range of the next folio, while the rest of that folio could still be dirty. So clearing dirty on the submitted sectors doesn't call folio_clear_dirty_for_io() and we have the same issue. Since extent_write_cache_pages() skips these batch submitted folios (they are already marked for writeback from submission by the preceding folio), we must add the extra write protection in lock_delalloc_folios(). 3. For inline extents this will subtly risk losing writes that happen after/while we copy the inline extent but before we clear dirty on the folio. 4. For folios spanning EOF, mmap could tamper with the zeroed bytes past EOF and cause them to be persisted where future faults would improperly see them instead of zeros. 5. Finally, for compressed extents, we risk modifying the folios while we work on compressing them which will result in corrupted compressed data. Specifically, in run_delalloc_compressed() we queue up work to do compress_file_range() in BTRFS_COMPRESSION_CHUNK_SIZE (512K) chunks which will call btrfs_folio_clamp_clear_dirty() on the range. For non-subpage, this will always clear the whole folio, safely. For subpage, we risk a partial clear here as well. In particular, imagine a 2M folio broken up into 512K chunks of work which might start compression work on one chunk before all the chunks compress_file_range() workers have gotten far enough to finish clearing all the dirty bitmaps of the folio and getting to folio_clear_dirty_for_io(). Large folios on the edges of submission ranges are similarly at risk to be only partly cleared. This particular gap was introduced by a second patch in the same series: commit a4ef54dbb576 ("btrfs: make extent_range_clear_dirty_for_io() to handle sector size < page size cases") We cannot simply restore the call to folio_clear_dirty_for_io() because that also drops the dirty flag off the folio which violates invariants introduced for large folios by commit 334509ce9d07 ("btrfs: use dirty flag to check if an ordered extent needs to be truncated") and results in failing to invalidate clean folios past i_size, resulting in deadlocks. Therefore, to fix it, leave the existing semantics w.r.t. the folio's dirty flag (to preserve the correct invalidate behavior) but ensure that the other aspect of folio_clear_dirty_for_io(), folio_mkclean(), is run on the folio when we lock it for writeback. Finally, to help prevent similar regressions in the future, add a debug warning that triggers at the known corruption sites if we have failed to write protect the folio. Assisted-by: LLM (debug, reproduce, research fix, review patch) Fixes: 095be159f3eb ("btrfs: unify folio dirty flag clearing") Fixes: a4ef54dbb576 ("btrfs: make extent_range_clear_dirty_for_io() to handle sector size < page size cases") Reviewed-by: Qu Wenruo <wqu@suse.com> Signed-off-by: Boris Burkov <boris@bur.io> Signed-off-by: David Sterba <dsterba@suse.com>
2026-08-07btrfs: fix a lockdep caused by path resolution during device scanQu Wenruo1-36/+1
[BUG] There is a lockdep report related to device scan: ====================================================== WARNING: possible circular locking dependency detected 7.2.0-20260712.rc2.git0.e3321fa3034d.300.fc44.s390x+debug #1 Not tainted ------------------------------------------------------ (udev-worker)/1653 is trying to acquire lock: 0000006919232220 (&type->i_mutex_dir_key#2){++++}-{3:3}, at: lookup_slow+0x3e/0x70 but task is already holding lock: 00000069238564d8 (&fs_devs->device_list_mutex){+.+.}-{3:3}, at: device_list_add.constprop.0+0x148/0xc60 which lock already depends on the new lock. the existing dependency chain (in reverse order) is: -> #5 (&fs_devs->device_list_mutex){+.+.}-{3:3}: lock_acquire+0x150/0x3f0 __mutex_lock+0xba/0xdc0 mutex_lock_nested+0x32/0x40 write_all_supers+0x7a/0x670 btrfs_sync_log+0xae6/0xdd0 btrfs_sync_file+0x4fa/0x7a0 __s390x_sys_fsync+0x52/0xa0 __do_syscall+0x172/0x750 system_call+0x72/0x90 -> #4 (&fs_info->tree_log_mutex){+.+.}-{3:3}: lock_acquire+0x150/0x3f0 __mutex_lock+0xba/0xdc0 mutex_lock_nested+0x32/0x40 btrfs_sync_log+0xaba/0xdd0 btrfs_sync_file+0x4fa/0x7a0 __s390x_sys_fsync+0x52/0xa0 __do_syscall+0x172/0x750 system_call+0x72/0x90 -> #3 (btrfs_trans_num_extwriters){.+.+}-{0:0}: lock_acquire+0x150/0x3f0 join_transaction+0x108/0x680 start_transaction+0x21a/0x660 btrfs_join_transaction+0x32/0x40 btrfs_dirty_inode+0x52/0xf0 touch_atime+0x90/0xc0 filemap_read+0x446/0x450 vfs_read+0x208/0x370 ksys_read+0x88/0x120 __do_syscall+0x172/0x750 system_call+0x72/0x90 -> #2 (btrfs_trans_num_writers){.+.+}-{0:0}: reacquire_held_locks+0x14c/0x240 __lock_release.isra.0+0xd8/0x380 lock_release+0xf6/0x270 percpu_up_read+0x28/0xf0 __btrfs_end_transaction+0x178/0x1f0 btrfs_dirty_inode+0x82/0xf0 touch_atime+0x90/0xc0 btrfs_file_mmap_prepare+0x8c/0xa0 __mmap_region+0x214/0x780 mmap_region+0x108/0x160 do_mmap+0x402/0x5a0 vm_mmap_pgoff+0x156/0x230 ksys_mmap_pgoff+0x17e/0x220 __s390x_sys_old_mmap+0xa8/0x140 __do_syscall+0x172/0x750 system_call+0x72/0x90 -> #1 (&mm->mmap_lock){++++}-{3:3}: lock_acquire+0x150/0x3f0 __might_fault+0x7a/0xa0 filldir64+0x11c/0x210 offset_readdir+0x92/0x200 iterate_dir+0xcc/0x2d0 __do_sys_getdents64+0x7a/0x130 __do_syscall+0x172/0x750 system_call+0x72/0x90 -> #0 (&type->i_mutex_dir_key#2){++++}-{3:3}: check_prev_add+0x160/0xf40 __lock_acquire+0x12aa/0x15a0 lock_acquire+0x150/0x3f0 down_read+0x5a/0x280 lookup_slow+0x3e/0x70 path_lookupat+0x1f0/0x370 filename_lookup+0xce/0x1f0 kern_path+0x48/0x70 is_same_device+0x146/0x300 device_list_add.constprop.0+0x1be/0xc60 btrfs_scan_one_device+0x13a/0x2f0 btrfs_control_ioctl+0x110/0x1e0 __s390x_sys_ioctl+0xfa/0x130 __do_syscall+0x172/0x750 system_call+0x72/0x90 other info that might help us debug this: Chain exists of: &type->i_mutex_dir_key#2 --> &fs_info->tree_log_mutex --> &fs_devs->device_list_mutex Possible unsafe locking scenario: CPU0 CPU1 ---- ---- lock(&fs_devs->device_list_mutex); lock(&fs_info->tree_log_mutex); lock(&fs_devs->device_list_mutex); rlock(&type->i_mutex_dir_key#2); *** DEADLOCK *** 2 locks held by (udev-worker)/1653: #0: 0000016c727051c8 (uuid_mutex){+.+.}-{3:3}, at: btrfs_control_ioctl+0x102/0x1e0 #1: 00000069238564d8 (&fs_devs->device_list_mutex){+.+.}-{3:3}, at: device_list_add.constprop.0+0x148/0xc60 stack backtrace: CPU: 2 UID: 0 PID: 1653 Comm: (udev-worker) Not tainted 7.2.0-20260712.rc2.git0.e3321fa3034d.300.fc44.s390x+debug #1 PREEMPT Hardware name: IBM 3931 A01 701 (LPAR) Call Trace: [<0000016c70680e3e>] dump_stack_lvl+0xae/0x108 [<0000016c7078aa44>] print_circular_bug+0x1a4/0x230 [<0000016c7078ac5c>] check_noncircular+0x18c/0x1b0 [<0000016c7078c030>] check_prev_add+0x160/0xf40 [<0000016c7078fbaa>] __lock_acquire+0x12aa/0x15a0 [<0000016c7078fff0>] lock_acquire+0x150/0x3f0 [<0000016c7180e2fa>] down_read+0x5a/0x280 [<0000016c70b88dde>] lookup_slow+0x3e/0x70 [<0000016c70b8f5d0>] path_lookupat+0x1f0/0x370 [<0000016c70b900ae>] filename_lookup+0xce/0x1f0 [<0000016c70b90218>] kern_path+0x48/0x70 [<0000016c70f5b4a6>] is_same_device+0x146/0x300 [<0000016c70f67cfe>] device_list_add.constprop.0+0x1be/0xc60 [<0000016c70f688da>] btrfs_scan_one_device+0x13a/0x2f0 [<0000016c70ed9cb0>] btrfs_control_ioctl+0x110/0x1e0 [<0000016c70b97d0a>] __s390x_sys_ioctl+0xfa/0x130 [<0000016c718004d2>] __do_syscall+0x172/0x750 [<0000016c718155d2>] system_call+0x72/0x90 [CAUSE] Btrfs device scan will call is_same_device() with device_list_mutex held. But is_same_device() will call kern_path() which will do path resolution and lock the inode. So device scan has the following lock sequence: mutex_lock(device_list_mutex) from device_list_add() | v inode_lock_shared() from lookup_slow() during kern_path(). Meanwhile another thread is fsyncing, which has the following lock sequence: inode_lock() from btrfs_inode_lock() inside btrfs_direct_write() | v mutex_lock(tree_log_mutex() from btrfs_sync_log(), which is further triggered from iomap_dio_complete()->generic_write_sync()->btrfs_sync_file(). | v mutex_lock(device_list_mutex) from write_all_supers() inside btrfs_sync_log(). So the device scan has a reversed lock sequence, compared to the fsync one, this means we can have the following deadlock: Device scan | Fsync ----------------------------------------+-------------------------------- device_list_mutex locked | | inode locked | try to lock device_list_mutex try to lock inode | [FIX] Instead of a full path lookup, use dev_t to determine if two device paths are pointing to the same block device. Inside kernel dev_t is going to uniquely determine a block device, and the device path lookup is already done by lookup_bdev(), which is done without device_list_mutex held, thus no reversed locking sequence. Reported-by: Christian Borntraeger <borntraeger@linux.ibm.com> Link: https://lore.kernel.org/linux-btrfs/5a9d9847-4ae6-43c4-afdc-6e5fa51d6117@linux.ibm.com/ Fixes: 2e8b6bc0ab41 ("btrfs: avoid unnecessary device path update for the same device") Tested-by: Christian Borntraeger <borntraeger@linux.ibm.com> Reviewed-by: Filipe Manana <fdmanana@suse.com> Signed-off-by: Qu Wenruo <wqu@suse.com> Signed-off-by: David Sterba <dsterba@suse.com>
2026-08-07btrfs: change block group reclaim_mark to boolSun YangKai3-3/+5
The reclaim_mark field in struct btrfs_block_group was a u64 that was incremented when marking block groups for reclaim during sweeping, but the actual counter value was never used - only the zero/non-zero state mattered for determining if a block group needed reclaim. Convert it to a bool to properly reflect its usage and reduce memory footprint by 8 bytes. Update assignments to use true/false instead of increment and zero. Reviewed-by: Boris Burkov <boris@bur.io> Signed-off-by: Sun YangKai <sunk67188@gmail.com> Reviewed-by: David Sterba <dsterba@suse.com> Signed-off-by: David Sterba <dsterba@suse.com>
2026-08-07btrfs: sink idmap parameter to __btrfs_ioctl_snap_create()David Sterba1-5/+3
The 'idmap' parameter is derived from 'file' that we also pass to __btrfs_ioctl_snap_create(), assign it inside the function. Reviewed-by: Qu Wenruo <wqu@suse.com> Signed-off-by: David Sterba <dsterba@suse.com>
2026-08-07btrfs: use kvmalloc() for stripe buffer of scrub_stripeQu Wenruo3-103/+68
Currently we're using scrub_stripe::folios[] to store all contents of a stripe. This means we need all the extra work to handle things like sub-page cases, and also require larger folios to handle bs > ps cases. On the other hand, it's not hard to allocate a 64K large folio to cover the full stripe, getting rid of the cross-page handling. Furthermore, even if that large folio allocation failed, we can still use vmalloc() to allocate a virtually contiguous space and still get rid of cross-page handling. This patch will go with kvmalloc() to allocate 64K of memory for the stripe buffer, thus getting rid of all the complex cross-page handling. The following aspects can be greatly simplified: - Checksum verification for both data and metadata No more per-page iteration, all in one go. - RAID56 data caching Just copy the buffer into the RAID56 pages. - No more kaddr/paddr grabbing For most cases the virtual address is enough for csum calculation and io submission. - Bio assembly There is already the helper bio_add_vmalloc() to queue vmallocated memory into a bio. Although it means we have something else to be concerned about: - Bio assembly If the memory is vmallocated, we need to use bio_add_vmalloc() Otherwise use the existing bio_add_page(). - Read endio For vmallocated memory, we need to call invalidate_kernel_vmap_range(). - Scrub bbio bvec size Since scrub_stripe::buffer is kvmallocated, we also need to enlarge the scrub bbio, to be able to handle the worst case, where all 64KiB is allocated by discontiguous 4K physical pages. Signed-off-by: Qu Wenruo <wqu@suse.com> Reviewed-by: David Sterba <dsterba@suse.com> Signed-off-by: David Sterba <dsterba@suse.com>
2026-08-07btrfs: scrub: implement calc_sector_number() in a faster wayQu Wenruo1-12/+13
Currently calc_sector_number() is implemented by comparing the first bvec of the bbio against all blocks inside a scrub_stripe. This implementation is a little inefficient, and depends on how the scrub buffer is implemented. One of the reason implementing such complex function is that, we do not save the original bvec_iter inside a write btrfs_bio. Although a read bbio has btrfs_bio::saved_iter to get the original logical bytenr, it's not implemented for write bios. On the other hand, since commit 81cea6cd7041 ("btrfs: remove btrfs_bio::fs_info by extracting it from btrfs_bio::inode"), we always set the btrfs_bio::file_offset as the logical bytenr for scrub, and that member will not be modified during IO. So this means we have a stable way to determine the logical bytenr for a scrub bio, now calc_sector_number() is just as simple as: return (bbio->file_offset - stripe->logical) >> sectorsize_bits; Since we're here, also add an ASSERT() to make sure the bbio is inside the stripe, and change the return type to unsigned int to be extra safe. Reviewed-by: Johannes Thumshirn <johannes.thumshirn@wdc.com> Signed-off-by: Qu Wenruo <wqu@suse.com> Reviewed-by: David Sterba <dsterba@suse.com> Signed-off-by: David Sterba <dsterba@suse.com>
2026-08-07btrfs: factor out common scrub read endio into a helperQu Wenruo1-22/+22
For both scrub_repair_read_endio() and scrub_read_endio(), they share the same bitmap update and bio put. Factor out the common code into a helper to reduce duplication. Reviewed-by: Johannes Thumshirn <johannes.thumshirn@wdc.com> Signed-off-by: Qu Wenruo <wqu@suse.com> Reviewed-by: David Sterba <dsterba@suse.com> Signed-off-by: David Sterba <dsterba@suse.com>
2026-08-07btrfs: remove SCRUB_MAX_SECTORS_PER_BLOCKQu Wenruo1-14/+0
The last user of this macro is removed in commit 001e3fc263ce ("btrfs: scrub: remove scrub_block and scrub_sector structures"). Now that macro is only utilized in an ASSERT(), which no longer makes much sense. Just remove it completely. Reviewed-by: Johannes Thumshirn <johannes.thumshirn@wdc.com> Signed-off-by: Qu Wenruo <wqu@suse.com> Reviewed-by: David Sterba <dsterba@suse.com> Signed-off-by: David Sterba <dsterba@suse.com>
2026-08-07btrfs: get rid of useless label in btrfs_create_dio_extent()Filipe Manana1-2/+1
There's no point in having a label where under it we do nothing but return a variable. So remove it and directly return where we used to goto. Reviewed-by: Johannes Thumshirn <johannes.thumshirn@wdc.com> Reviewed-by: Qu Wenruo <wqu@suse.com> Signed-off-by: Filipe Manana <fdmanana@suse.com> Reviewed-by: David Sterba <dsterba@suse.com> Signed-off-by: David Sterba <dsterba@suse.com>
2026-08-07btrfs: fix extent map leak in NOCOW direct I/O writeShuangpeng Bai1-6/+13
btrfs_dio_iomap_begin() calls btrfs_get_extent(), which returns an extent map reference that must be dropped on all exit paths. For direct writes into a NOCOW range, btrfs_get_blocks_direct_write() keeps using that extent map and asks btrfs_create_dio_extent() to allocate the ordered extent. If that fails, for example because btrfs_alloc_ordered_extent() fails, the function returns the error without dropping the input extent map. The PREALLOC path avoided this by dropping the input extent map before replacing it with the newly created one. Check the error from btrfs_create_dio_extent() before replacing the map and drop the input extent map on failure. Fixes: 5f9a8a51d8b9 ("Btrfs: add semaphore to synchronize direct IO writes with fsync") CC: stable@vger.kernel.org Reviewed-by: Qu Wenruo <wqu@suse.com> Reviewed-by: Filipe Manana <fdmanana@suse.com> Signed-off-by: Shuangpeng Bai <shuangpeng.kernel@gmail.com> Signed-off-by: Filipe Manana <fdmanana@suse.com> Signed-off-by: David Sterba <dsterba@suse.com>
2026-08-07btrfs: remove "usebackuproot" mount optionQu Wenruo1-11/+0
This mount option is marked deprecated since the introduction of "rescue=" mount option group, in v5.9. That's already a long time ago, and it should be safe to completely remove the old "usebackuproot" mount option now. Reviewed-by: Johannes Thumshirn <johannes.thumshirn@wdc.com> Reviewed-by: Neal Gompa <neal@gompa.dev> Signed-off-by: Qu Wenruo <wqu@suse.com> Signed-off-by: David Sterba <dsterba@suse.com>
2026-08-07btrfs: add "rescue=usebackuproot" into "rescue=all" shortcutQu Wenruo1-0/+1
The mount option "rescue=all" should be a shortcut to include all "rescue=" mount options. But unfortunately "rescue=usebackuproot" is not included. Include that option so "rescue=all" has a better chance to mount a corrupted fs. Reviewed-by: Johannes Thumshirn <johannes.thumshirn@wdc.com> Reviewed-by: Neal Gompa <neal@gompa.dev> Signed-off-by: Qu Wenruo <wqu@suse.com> Signed-off-by: David Sterba <dsterba@suse.com>
2026-08-07btrfs: add "rescue=usebackuproot" into forced read-only optionsQu Wenruo2-3/+4
According to btrfs(5) man page, all rescue options should require a read-only mount. But that read-only check is only introduced for newer rescue options, not for the pre-existing "usebackuproot" one. Furthermore, a filesystem that requires "rescue=" mount option already means it's corrupted, even if "rescue=usebackuproot" allowed the fs to be mounted RW, one should not trust such fs anymore until a comprehensive btrfs-check run and proper evaluation. Change the behavior to match the document, and since "rescue=usebackuproot" is now a full RO mount option, it is no longer a one-shot option, therefore remove it from btrfs_clear_oneshot_options(). Reviewed-by: Johannes Thumshirn <johannes.thumshirn@wdc.com> Reviewed-by: Neal Gompa <neal@gompa.dev> Signed-off-by: Qu Wenruo <wqu@suse.com> Signed-off-by: David Sterba <dsterba@suse.com>