| Age | Commit message (Collapse) | Author | Files | Lines |
|
The 'tree_id' parameter in btrfs_search_path_in_tree() was only being
used in order to fetch the root tree to be considered for the
search. For this same reason this function was also requiring a 'struct
btrfs_fs_info' parameter. This commit replaces these two parameters with
a single 'struct btrfs_root' one, which identifies from which root tree
the search should happen.
This function only has one caller, the inode lookup ioctl, which knows
how to provide the root tree for each case. In fact, if args->treeid ==
0, then we don't even have to allocate a new root tree object, and we
can reuse the one provided by the ioctl system call, thus avoiding an
extra allocation.
Signed-off-by: Miquel Sabaté Solà <mssola@mssola.com>
Reviewed-by: Filipe Manana <fdmanana@suse.com>
Signed-off-by: Filipe Manana <fdmanana@suse.com>
Signed-off-by: David Sterba <dsterba@suse.com>
|
|
Previously btrfs forces direct writes to fall back to buffered ones if the
inode has data checksum or the profile has duplication.
That fallback is to avoid the content being modified that the final
content may mismatch with the checksum or the other mirrors.
That brings a pretty huge performance cost, which already caused some
concern at that time.
But later upstream commit c9d114846b38 ("iomap: add a flag to bounce
buffer direct I/O") introduced a new method by copying the content into
new pages, and do all the operations based on the newly allocated pages.
So let btrfs to utilize the new flag for direct writes if we require
stable folios.
There is a quick benchmark, using the following fio setup:
fio --name=randwrite --filename $mnt/foobar --ioengine=libaio --size=4G \
--rw=randwrite --iodepth=64 --runtime=60 --time_based --direct=1 \
--bs=$blocksize
Unit is MiB/s.
Blocksize | Zero-copy (*) | Buffered | Bounce
-----------+---------------+----------+-----------
4K | 35.1 | 17.1 | 33.8
64K | 522 | 251 | 492
*: This is done by reverting the commit 968f19c5b1b7 ("btrfs: always
fallback to buffered write if the inode requires checksum")
Although with page bouncing the performance is only around 95% of
true-zero copy, it's still almost double the performance of buffered
fallback.
There will be a small change in behavior, since we're using
IOMAP_DIO_BOUNCE flag to allocate new folios, NOWAIT flag will
immediately fail.
So for true NOWAIT direct IOs, NODATASUM and RAID0/SINGLE profiles are
still required.
Signed-off-by: Qu Wenruo <wqu@suse.com>
Signed-off-by: David Sterba <dsterba@suse.com>
|
|
[BUG]
Syzbot reported a bug that there can be conflicting OEs for the same
range:
BTRFS critical (device loop4): panic in insert_ordered_extent:264: overlapping ordered extents, existing oe file_offset 16384 num_bytes 430080 flags 0x1089, new oe file_offset 16384 num_bytes 430080 flags 0x80 (errno=-17 Object alrea[ 179.162726][ T6897] BTRFS critical (device loop4): panic in insert_ordered_extent:264: overlapping ordered extents, existing oe file_offset 16384 num_bytes 430080 flags 0x1089, new oe file_offset 16384 num_bytes 430080 flags 0x80 (errno=-17 Object already exists)
------------[ cut here ]------------
kernel BUG at fs/btrfs/ordered-data.c:264!
Hardware name: Google Google Compute Engine/Google Compute Engine, BIOS Google 05/09/2026
RIP: 0010:btrfs_alloc_ordered_extent+0x943/0xad0
Call Trace:
<TASK>
cow_file_range+0x744/0x12a0
fallback_to_cow+0x5ea/0xa00
run_delalloc_nocow+0x110c/0x17a0
btrfs_run_delalloc_range+0xbe4/0x1c20
writepage_delalloc+0x104d/0x1ba0
btrfs_writepages+0x1667/0x28b0
do_writepages+0x338/0x560
filemap_fdatawrite_range+0x1f2/0x300
btrfs_fdatawrite_range+0x54/0xf0
btrfs_direct_write+0x6a0/0xc30
btrfs_do_write_iter+0x329/0x790
do_iter_readv_writev+0x624/0x8d0
vfs_writev+0x34c/0x990
__se_sys_pwritev2+0x17a/0x2a0
do_syscall_64+0x174/0x580
entry_SYSCALL_64_after_hwframe+0x77/0x7f
</TASK>
---[ end trace 0000000000000000 ]---
[CAUSE]
Since commit ff66fe666233 ("btrfs: fix incorrect buffered IO fallback
for append direct writes"), if the direct IO finished short, we will
revert the isize back to the original one, so that append writes can be
respected during the buffered fallback.
Normally we rely on lock_and_cleanup_extent_if_need() function during
buffered writeback to wait for any existing ordered extents.
But that ordered extent waiting only happens if the start_pos is inside
the isize.
Since we have reverted the isize during failed direct IO, we will not
wait for any ordered extents.
This means we can have a race where the direct IO OE is still in the
tree, finished but not yet removed, then we're inserting the OE for the
buffered write, causing the above crash.
[FIX]
Make the OE wait to be unconditional, to handle the reverted isize
situation.
And since lock_and_cleanup_extent_if_need() now either lock the
extents or return -EAGAIN, also remove the branches that handles
no-extent-locked cases, and rename it to remove the "_if_need" suffix.
The following micro benchmark shows the runtime difference for
btrfs_buffered_write(), doing `xfs_io -f -c "pwrite 0 1m"` workload,
all values are the average runtime in nano seconds.
function runtime | before | after
-----------------------------------+-------------+---------------
lock_and_cleanup_extent_if_need() | 58.2 | 183.0
btrfs_buffered_write() | 2115.6 | 2973.3
The overall runtime of btrfs_buffered_write() is still pretty
tiny (still less than 3 micro seconds), I'd say the extra cost is still
acceptable.
An alternative to fix this problem is to wait ordered extents during
iomap_end() where the isize revert is done.
But that solution will break nowait requirement, as if a nowait direct
IO finished short, we have to wait for the OEs unconditionally or the
next append buffered IO can still hit the same problem.
So here we have to move the wait cost to buffered write, but at least
the code is slightly more streamline.
Reported-by: syzbot+ba2afde329fc27e3f22e@syzkaller.appspotmail.com
Link: https://syzkaller.appspot.com/bug?extid=ba2afde329fc27e3f22e
Fixes: ff66fe666233 ("btrfs: fix incorrect buffered IO fallback for append direct writes")
Reviewed-by: Filipe Manana <fdmanana@suse.com>
Signed-off-by: Qu Wenruo <wqu@suse.com>
Signed-off-by: David Sterba <dsterba@suse.com>
|
|
There's no need to call list_del_init() against each entry when freeing
the list, as the list is local and we are freeing the entry.
Reviewed-by: Qu Wenruo <wqu@suse.com>
Signed-off-by: Filipe Manana <fdmanana@suse.com>
Signed-off-by: David Sterba <dsterba@suse.com>
|
|
When freeing the entries from the list there is no need to initialize
the list member in an entry, since we are immediately freeing it. So use
simple list_del() instead of list_del_init().
Reviewed-by: Qu Wenruo <wqu@suse.com>
Signed-off-by: Filipe Manana <fdmanana@suse.com>
Signed-off-by: David Sterba <dsterba@suse.com>
|
|
Use AUTO_KFREE() for the folios array, avoiding two kfree() calls, one of
them in a very specific error path.
Reviewed-by: Qu Wenruo <wqu@suse.com>
Signed-off-by: Filipe Manana <fdmanana@suse.com>
Signed-off-by: David Sterba <dsterba@suse.com>
|
|
There's no need to have one list for each loop to defrag each subrange and
then another one to free each subrange (struct defrag_target_range).
We can do it in a single loop, freeing each subrange after defragging,
plus no need to delete each subrange from the list since we immediately
free it.
Reviewed-by: Qu Wenruo <wqu@suse.com>
Signed-off-by: Filipe Manana <fdmanana@suse.com>
Reviewed-by: David Sterba <dsterba@suse.com>
Signed-off-by: David Sterba <dsterba@suse.com>
|
|
Syzbot reported the following warning recently:
[157.672][ T6611] BTRFS info (device loop0): turning on flush-on-commit
[157.672][ T6611] BTRFS info (device loop0): enabling free space tree
[157.672][ T6611] BTRFS info (device loop0): enabling auto defrag
[157.672][ T6611] BTRFS info (device loop0): use lzo compression, level 1
[157.672][ T6611] BTRFS info (device loop0): max_inline set to 4096
[158.094][ T5608] BTRFS info (device loop2): last unmount of filesystem c9fe44da-de57-406a-8241-57ec7d4412cf
[160.073][ T6656] BTRFS info (device loop0 state M): max_inline set to 4096
[160.418][ T5611] BTRFS info (device loop0): last unmount of filesystem ab8108e1-bea5-4a9f-94c9-a3ff208d732a
[160.432][ T6662] loop2: detected capacity change from 0 to 32768
[160.438][ T6662] BTRFS: device fsid c9fe44da-de57-406a-8241-57ec7d4412cf devid 1 transid 8 /dev/loop2 (7:2) scanned by syz.2.74 (6662)
[160.459][ T6662] BTRFS info (device loop2): first mount of filesystem c9fe44da-de57-406a-8241-57ec7d4412cf
[160.459][ T6662] BTRFS info (device loop2): using crc32c checksum algorithm
[160.634][ T1187] ------------[ cut here ]------------
[160.634][ T1187] test_bit(BTRFS_FS_STATE_NO_DELAYED_IPUT, &fs_info->fs_state)
[160.634][ T1187] WARNING: fs/btrfs/inode.c:3596 at btrfs_add_delayed_iput+0x2e3/0x340, CPU#0: kworker/u8:10/1187
[160.634][ T1187] Modules linked in:
[160.634][ T1187] CPU: 0 UID: 0 PID: 1187 Comm: kworker/u8:10 Not tainted syzkaller #0 PREEMPT_{RT,(full)}
[160.634][ T1187] Hardware name: Google Google Compute Engine/Google Compute Engine, BIOS Google 04/18/2026
[160.634][ T1187] Workqueue: btrfs-endio-write btrfs_work_helper
[160.634][ T1187] RIP: 0010:btrfs_add_delayed_iput+0x2e3/0x340
[160.634][ T1187] Code: 53 a3 45 (...)
[160.634][ T1187] RSP: 0018:ffffc900065d77c8 EFLAGS: 00010293
[160.634][ T1187] RAX: ffffffff83e5f502 RBX: ffff88805aba0000 RCX: ffff888029768000
[160.634][ T1187] RDX: 0000000000000000 RSI: 0000000000000000 RDI: 0000000000000000
[160.634][ T1187] RBP: dffffc0000000000 R08: 0000000000000000 R09: 0000000000000000
[160.634][ T1187] R10: dffffc0000000000 R11: ffffed100b574497 R12: 0000000000000001
[160.634][ T1187] R13: dffffc0000000000 R14: ffff888061194788 R15: 0000000000000200
[160.634][ T1187] FS: 0000000000000000(0000) GS:ffff888126186000(0000) knlGS:0000000000000000
[160.634][ T1187] CS: 0010 DS: 0000 ES: 0000 CR0: 0000000080050033
[160.634][ T1187] CR2: 00007fe553a3f000 CR3: 00000000596c2000 CR4: 00000000003526f0
[160.634][ T1187] Call Trace:
[160.634][ T1187] <TASK>
[160.634][ T1187] btrfs_put_ordered_extent+0x18f/0x430
[160.634][ T1187] btrfs_finish_one_ordered+0xf63/0x2680
[160.634][ T1187] ? __pfx_btrfs_finish_one_ordered+0x10/0x10
[160.634][ T1187] ? do_raw_spin_lock+0x12b/0x2f0
[160.634][ T1187] ? lock_acquire+0x106/0x350
[160.634][ T1187] ? __pfx_do_raw_spin_lock+0x10/0x10
[160.634][ T1187] btrfs_work_helper+0x38b/0xc20
[160.634][ T1187] ? process_scheduled_works+0xa70/0x1860
[160.634][ T1187] process_scheduled_works+0xb5d/0x1860
[160.634][ T1187] ? __pfx_process_scheduled_works+0x10/0x10
[160.634][ T1187] ? assign_work+0x3d5/0x5e0
[160.634][ T1187] worker_thread+0xa53/0xfc0
[160.634][ T1187] kthread+0x388/0x470
[160.634][ T1187] ? __pfx_worker_thread+0x10/0x10
[160.635][ T1187] ? __pfx_kthread+0x10/0x10
[160.635][ T1187] ret_from_fork+0x514/0xb70
[160.635][ T1187] ? __pfx_ret_from_fork+0x10/0x10
[160.635][ T1187] ? __switch_to+0xc79/0x1410
[160.635][ T1187] ? __pfx_kthread+0x10/0x10
[160.635][ T1187] ret_from_fork_asm+0x1a/0x30
[160.635][ T1187] </TASK>
[160.635][ T1187] Kernel panic - not syncing: kernel: panic_on_warn set ...
It means we add a delayed iput created after we last ran delayed iputs in
close_ctree() and set the flag BTRFS_FS_STATE_NO_DELAYED_IPUT in fs_info.
This happens when using autodefrag and more likely to happen if we use
flushoncommit too. The steps are the following:
1) Unmount starts, all delalloc is flushed and we enter close_ctree();
2) In close_ctree() we park the cleaner kthread, but while we wait for it
to park, it's in:
btrfs_run_defrag_inodes()
btrfs_run_defrag_inode()
btrfs_defrag_file()
defrag_one_cluster()
defrag_one_range()
defrag_one_locked_target()
And dirties some folios from an inode;
3) The cleaner kthread parks and we proceed in close_ctree(), waiting
for all ordered extents, running delayed iputs and setting the flag
BTRFS_FS_STATE_NO_DELAYED_IPUT in fs_info;
4) Later in close_ctree() we call btrfs_commit_super(), which commits the
current transaction. Because we are mounted with flushoncommit, the
transaction commit flushes delalloc and waits for the resulting ordered
extent to complete;
5) The ordered extents from the flushed delalloc created by autodefrag
complete and create delayed iputs, triggering the warning:
WARN_ON_ONCE(test_bit(BTRFS_FS_STATE_NO_DELAYED_IPUT, &fs_info->fs_state));
in btrfs_add_delayed_iput()
6) Further below in close_ctree() we will hit the following assertion:
ASSERT(list_empty(&fs_info->delayed_iputs));
Since we don't expect any more delayed iputs.
Fix this by flushing delalloc and waiting for the ordered extents right
after we parked the cleaner kthread and waiting for autodefrag in
close_ctree().
Reported-by: syzbot+6a843bf8604711c8fab0@syzkaller.appspotmail.com
Link: https://lore.kernel.org/linux-btrfs/6a1ee507.b4221f80.1326c5.0004.GAE@google.com/
Reviewed-by: Qu Wenruo <wqu@suse.com>
Signed-off-by: Filipe Manana <fdmanana@suse.com>
Signed-off-by: David Sterba <dsterba@suse.com>
|
|
While running fsstress with autodefrag and flushoncommit, hit a deadlock
due to the fact that defrag reserves delalloc space while it's holding
dirty and locked folios, besides the extent range lock. The stack traces
are the following:
[958.624] task:kworker/u50:3 state:D stack:0 pid:20365 tgid:20365 ppid:2 task_flags:0x4208060 flags:0x00080000
[958.626] Workqueue: events_unbound btrfs_async_reclaim_metadata_space [btrfs]
[958.627] Call Trace:
[958.628] <TASK>
[958.628] __schedule+0x4be/0x10f0
[958.629] ? preempt_count_add+0x69/0xa0
[958.630] schedule+0x26/0xd0
[958.631] wait_current_trans+0x102/0x160 [btrfs]
[958.632] ? __pfx_autoremove_wake_function+0x10/0x10
[958.633] start_transaction+0x374/0x900 [btrfs]
[958.634] btrfs_commit_current_transaction+0x1d/0x70 [btrfs]
[958.635] flush_space+0xca/0x5e0 [btrfs]
[958.636] ? _raw_spin_unlock+0x15/0x30
[958.637] ? btrfs_reduce_alloc_profile+0x8c/0x190 [btrfs]
[958.639] ? _raw_spin_unlock+0x15/0x30
[958.640] ? calc_available_free_space.isra.0+0x6f/0x110 [btrfs]
[958.641] do_async_reclaim_metadata_space+0x84/0x190 [btrfs]
[958.642] btrfs_async_reclaim_metadata_space+0x64/0x80 [btrfs]
[958.644] process_one_work+0x19d/0x3a0
[958.644] worker_thread+0x1c4/0x330
[958.645] ? __pfx_worker_thread+0x10/0x10
[958.646] kthread+0xfc/0x130
[958.647] ? __pfx_kthread+0x10/0x10
[958.648] ret_from_fork+0x1f7/0x2c0
[958.648] ? __pfx_kthread+0x10/0x10
[958.649] ret_from_fork_asm+0x1a/0x30
[958.650] </TASK>
[958.651] task:kworker/u49:7 state:D stack:0 pid:52990 tgid:52990 ppid:2 task_flags:0x4208060 flags:0x00080000
[958.653] Workqueue: writeback wb_workfn (flush-btrfs-334)
[958.655] Call Trace:
[958.655] <TASK>
[958.656] __schedule+0x4be/0x10f0
[958.657] ? __blk_flush_plug+0xe9/0x140
[958.658] schedule+0x26/0xd0
[958.658] io_schedule+0x42/0x70
[958.659] folio_wait_bit_common+0x12b/0x330
[958.660] ? folio_wait_bit_common+0x100/0x330
[958.662] ? __pfx_wake_page_function+0x10/0x10
[958.663] extent_write_cache_pages+0x599/0x830 [btrfs]
[958.664] ? acpi_fwnode_get_reference_args+0x1fa/0x270
[958.665] btrfs_writepages+0x77/0x130 [btrfs]
[958.666] ? __pfx_end_bbio_data_write+0x10/0x10 [btrfs]
[958.667] do_writepages+0xc6/0x160
[958.668] __writeback_single_inode+0x42/0x310
[958.669] writeback_sb_inodes+0x231/0x570
[958.670] wb_writeback+0x8a/0x340
[958.671] wb_workfn+0xbf/0x450
[958.672] ? finish_task_switch.isra.0+0xc1/0x350
[958.673] process_one_work+0x19d/0x3a0
[958.673] worker_thread+0x1c4/0x330
[958.674] ? __pfx_worker_thread+0x10/0x10
[958.675] kthread+0xfc/0x130
[958.676] ? __pfx_kthread+0x10/0x10
[958.676] ret_from_fork+0x1f7/0x2c0
[958.677] ? __pfx_kthread+0x10/0x10
[958.678] ret_from_fork_asm+0x1a/0x30
[958.679] </TASK>
[958.679] task:btrfs-cleaner state:D stack:0 pid:296750 tgid:296750 ppid:2 task_flags:0x208040 flags:0x00080000
[958.681] Call Trace:
[958.682] <TASK>
[958.682] __schedule+0x4be/0x10f0
[958.683] schedule+0x26/0xd0
[958.684] handle_reserve_ticket+0x1b9/0x2c0 [btrfs]
[958.685] ? __pfx_autoremove_wake_function+0x10/0x10
[958.686] reserve_bytes+0x283/0x4c0 [btrfs]
[958.687] btrfs_reserve_metadata_bytes+0x18/0xb0 [btrfs]
[958.688] btrfs_delalloc_reserve_metadata+0x121/0x320 [btrfs]
[958.690] btrfs_delalloc_reserve_space+0x46/0xb0 [btrfs]
[958.691] btrfs_defrag_file+0x903/0x1110 [btrfs]
[958.692] btrfs_run_defrag_inodes+0x334/0x430 [btrfs]
[958.694] cleaner_kthread+0x97/0x1c0 [btrfs]
[958.694] ? __pfx_cleaner_kthread+0x10/0x10 [btrfs]
[958.696] kthread+0xfc/0x130
[958.696] ? __pfx_kthread+0x10/0x10
[958.697] ret_from_fork+0x1f7/0x2c0
[958.698] ? __pfx_kthread+0x10/0x10
[958.699] ret_from_fork_asm+0x1a/0x30
[958.700] </TASK>
[958.716] task:fsstress state:D stack:0 pid:296769 tgid:296769 ppid:296768 task_flags:0x400140 flags:0x00080000
[958.718] Call Trace:
[958.719] <TASK>
[958.719] __schedule+0x4be/0x10f0
[958.720] ? preempt_count_add+0x69/0xa0
[958.721] schedule+0x26/0xd0
[958.722] wb_wait_for_completion+0x79/0xc0
[958.723] ? __pfx_autoremove_wake_function+0x10/0x10
[958.724] __writeback_inodes_sb_nr+0xc5/0xf0
[958.725] try_to_writeback_inodes_sb+0x55/0x70
[958.726] btrfs_commit_transaction+0x19d/0xeb0 [btrfs]
[958.727] ? start_transaction+0x343/0x900 [btrfs]
[958.728] btrfs_mksubvol+0x28b/0x4e0 [btrfs]
[958.729] btrfs_mksnapshot+0x74/0xa0 [btrfs]
[958.730] __btrfs_ioctl_snap_create+0x194/0x210 [btrfs]
[958.732] btrfs_ioctl_snap_create_v2+0xef/0x150 [btrfs]
[958.733] btrfs_ioctl+0x7ec/0x2a70 [btrfs]
[958.734] ? __virt_addr_valid+0xe4/0x180
[958.735] ? __check_object_size+0x1cd/0x1f0
[958.736] ? kmem_cache_free+0x146/0x380
[958.737] ? _raw_spin_unlock+0x15/0x30
[958.738] ? do_sys_openat2+0x83/0xd0
[958.739] __x64_sys_ioctl+0x92/0xe0
[958.740] do_syscall_64+0x60/0x590
[958.741] ? clear_bhb_loop+0x60/0xb0
[958.742] entry_SYSCALL_64_after_hwframe+0x76/0x7e
[958.743] RIP: 0033:0x7f4431e108db
[958.744] RSP: 002b:00007ffcd147db20 EFLAGS: 00000246 ORIG_RAX: 0000000000000010
[958.746] RAX: ffffffffffffffda RBX: 0000000000000004 RCX: 00007f4431e108db
[958.747] RDX: 00007ffcd147eb90 RSI: 0000000050009417 RDI: 0000000000000005
[958.749] RBP: 0000000000000000 R08: 0000000000000000 R09: 0000000000000000
[958.751] R10: 0000000000000000 R11: 0000000000000246 R12: 00007ffcd147fbf0
[958.752] R13: 00007ffcd147eb90 R14: 0000000000000005 R15: 0000000000000003
[958.754] </TASK>
What happens is the following:
1) The cleaner kthread is running autodefrag, and in defrag_one_range()
it acquired all the folios for the range and locked them.
Then it locked the extent range in the inode's iotree.
It got two subranges from defrag_collect_targets(), the first one
with folio A and the second one with folio B.
After it defragged the first subrange, folio A remains locked and
dirty - it's only unlocked when defrag_one_range() returns.
When it attempts to defrag the second subrange (containing folio B),
btrfs_delalloc_reserve_space() creates a space reservation ticket,
due to lack of free metadata space and blocks waiting for the async
metadata reclaim task to free space and wake it up;
2) The async reclaim metadata task attempts to commit the current
transaction, but it blocks because there is another task that
started the commit first;
3) A task creating a snapshot is committing the transaction and
because the fs was mounted with flushoncommit, it calls
try_to_writeback_inodes_sb(), which spawns a task to flush
delalloc and waits for it to complete;
4) The task flushing delalloc (kworker/u49:7), finds that folio A for
the inode being defragged is dirty, so it tries to lock it...
But it blocks because folio A is locked by the defrag task (the
cleaner kthread) which is blocked waiting for the reservation
ticket to be served, but the async reclaim metadata task is
blocked waiting for the transaction commit, which in turn is
blocked waiting for the delalloc flush task, which is trying to
lock folio A, resulting in a deadlock.
The same type of problem can happen if the async reclaim task starts to
flush delalloc, as that requires both locking the folio and the extent
range in the inode's io tree, and in this case we don't need the fs to
be mounted with flushoncommit. This type of problem has ocurred several
times in the past with reflinks for example, where we had a dirty folio
while holding the extent range locked and then starting a transaction
blocked waiting for the async reclaim task due to lack of free metadata
space.
So fix this by reserving delalloc space before locking folios and locking
the extent range in the inode's iotree. We can not simply unlock the
folios for each subrange given by defrag_collect_targets() after we defrag
it because the same folio may be present too in the next subrange (due to
large folios).
Fixes: 22b398eeeed4 ("btrfs: defrag: introduce helper to defrag a contiguous prepared range")
Reviewed-by: Qu Wenruo <wqu@suse.com>
Signed-off-by: Filipe Manana <fdmanana@suse.com>
Signed-off-by: David Sterba <dsterba@suse.com>
|
|
Btrfs does not support variable stripe length yet, all RAID0/5/6/10
chunks have the fixed stripe length 64K for now.
Furthermore, btrfs_fs_info::stripesize is not the real chunk stripe
length, it's always the same value as sectorsize.
Remove btrfs_fs_info::stripesize, and for the only callsite utilizing
that member, replace it with fs_info->sectorsize instead.
Signed-off-by: Qu Wenruo <wqu@suse.com>
Reviewed-by: David Sterba <dsterba@suse.com>
Signed-off-by: David Sterba <dsterba@suse.com>
|
|
The nodesize and sectorsize are all u32 values, there is no need to use
u64 for local usage.
Furthermore some call sites also use "blocksize" or "bs" for sectorsize,
also change them to use the minimal type u32 instead.
Reviewed-by: Boris Burkov <boris@bur.io>
Signed-off-by: Qu Wenruo <wqu@suse.com>
Reviewed-by: David Sterba <dsterba@suse.com>
Signed-off-by: David Sterba <dsterba@suse.com>
|
|
prefixes
In case the current inode's path is a prefix of the given path, the helper
is_current_inode_path() will return true, which causes the single caller
to reset the current inode's path. While this is not a functional issue,
it makes the caller recompute the current inode's path later. It could
also become a problem in the future in case get new callers for
is_current_inode_path() in more sensitive contexts.
Example: the current inode path is "/foo/bar" and the path we compare
against is "/foo/bar_xyz".
Fix this by returning true only if we have exact matches.
Reviewed-by: Johannes Thumshirn <johannes.thumshirn@wdc.com>
Reviewed-by: Daniel Vacek <neelx@suse.com>
Signed-off-by: Filipe Manana <fdmanana@suse.com>
Signed-off-by: David Sterba <dsterba@suse.com>
|
|
The comment is wrong, because it's not about storing the ID of new
directories that were already created, instead it's about storing utimes
values for directories (both new and existing). The comment is wrong
because it was copy pasted from SEND_MAX_DIR_CREATED_CACHE_SIZE, but
forgot to update it afterwards.
Reviewed-by: Johannes Thumshirn <johannes.thumshirn@wdc.com>
Reviewed-by: Daniel Vacek <neelx@suse.com>
Signed-off-by: Filipe Manana <fdmanana@suse.com>
Signed-off-by: David Sterba <dsterba@suse.com>
|
|
On a zoned FS, btrfs_delayed_refs_rsv_refill() returns -EAGAIN whenever
the over-committed metadata plus the zone_unusable bytes exceeds the
usable size in a metadata block-group to avoid heavy over-commit of
metadata and early ENOSPC in one transaction.
If this happens while doing reclaim, the transaction is getting aborted.
Treat -EAGAIN as a soft, retryable condition in case of block-group
reclaim.
Reported-by: Damien Le Moal <dlemoal@kernel.org>
Fixes: 7bcb04de982f ("btrfs: zoned: cap delayed refs metadata reservation to avoid overcommit")
Reviewed-by: Filipe Manana <fdmanana@suse.com>
Signed-off-by: Johannes Thumshirn <johannes.thumshirn@wdc.com>
Signed-off-by: David Sterba <dsterba@suse.com>
|
|
Since v5.15 btrfs has support for block size < page size, but we still
only support 4K block size, while there is no special reason that we
cannot support 8K/16K/32K block sizes for 64K page size.
That 4K limit is completely arbitrary, and mostly to reduce test runtime
so we do not need to test all the extra block size combinations.
However that also limits the user choices, some users may understand
what they are doing, and want larger block sizes. In that case, fixed
4K block size for subpage routine is blocking our way.
Just remove that fixed 4K requirement for block size < page size.
This should not affect regular end users, since mkfs is already using 4K
block size as default for quite a while, and the existing bs == ps support is
always there.
But for power users, this allows extra block size support, and may
provide extra test coverage.
Reviewed-by: Johannes Thumshirn <johannes.thumshirn@wdc.com>
Signed-off-by: Qu Wenruo <wqu@suse.com>
Reviewed-by: David Sterba <dsterba@suse.com>
Signed-off-by: David Sterba <dsterba@suse.com>
|
|
Since commit bac3c2910c0c ("btrfs: remove 2K block size support") there
is no 2K block size support inside btrfs anymore.
Remove the stale comments of btrfs_supported_blocksize().
Reviewed-by: Johannes Thumshirn <johannes.thumshirn@wdc.com>
Signed-off-by: Qu Wenruo <wqu@suse.com>
Reviewed-by: David Sterba <dsterba@suse.com>
Signed-off-by: David Sterba <dsterba@suse.com>
|
|
V2 space cache has been the default mkfs option since btrfs-progs v5.15,
and commit 1e7bec1f7d65 ("btrfs: emit a warning about space cache v1
being deprecated") has already added a warning to show v1 space cache
has been deprecated.
It has been long enough that we should remove v1 space cache completely.
As the first step, disable v1 space cache by:
- Make "space_cache" mount option fallback to "nospace_cache"
- Make "space_cache=v1" fall back to "nospace_cache"
This is safer than forcing "space_cache=v2", as forcing v2 cache
requires removal of v1 cache and regenerating v2 cache.
Such operation can be slow, and takes extra metadata space, thus
it is not always safe for existing filesystems.
With this done, v1 cache mount will always fallback to nospace cache,
and mount option will not be able to force v1 space cache usage.
For example, even for a fs with v1 cache:
# btrfs ins dump-super test.img
superblock: bytenr=65536, device=test.img
---------------------------------------------------------
csum_type 0 (crc32c)
csum_size 4
csum 0xdce44b2c [match]
bytenr 65536
flags 0x1
( WRITTEN )
magic _BHRfS_M [match]
fsid 7d7c3bba-8211-4206-868d-10eedd5703f8
metadata_uuid 00000000-0000-0000-0000-000000000000
label
generation 9
root 30605312
[...]
compat_ro_flags 0x0 <<< No FST feature
incompat_flags 0x361
( MIXED_BACKREF |
BIG_METADATA |
EXTENDED_IREF |
SKINNY_METADATA |
NO_HOLES )
cache_generation 9 <<< Matches generation
uuid_tree_generation 9
Attempting to mount it will lead to no space cache other than v1 space cache:
# mount test.img /mnt/btrfs
# dmesg -t | tail -n 5
BTRFS: device fsid 7d7c3bba-8211-4206-868d-10eedd5703f8 devid 1 transid 9 /dev/loop0 (7:0) scanned by mount (1264)
BTRFS info (device loop0): first mount of filesystem 7d7c3bba-8211-4206-868d-10eedd5703f8
BTRFS info (device loop0): using crc32c checksum algorithm
BTRFS info (device loop0): turning on async discard
BTRFS info (device loop0): last unmount of filesystem 7d7c3bba-8211-4206-868d-10eedd5703f8
Even forcing v1 cache will not work, but fallback to the usual
nospace_cache:
# mount test.img -o space_cache=v1 /mnt/btrfs
# dmesg -t | tail -n 6
BTRFS warning: v1 space cache is deprecated, fallback to no space cache
BTRFS: device fsid 7d7c3bba-8211-4206-868d-10eedd5703f8 devid 1 transid 9 /dev/loop0 (7:0) scanned by mount (1264)
BTRFS info (device loop0): first mount of filesystem 7d7c3bba-8211-4206-868d-10eedd5703f8
BTRFS info (device loop0): using crc32c checksum algorithm
BTRFS info (device loop0): turning on async discard
BTRFS info (device loop0): last unmount of filesystem 7d7c3bba-8211-4206-868d-10eedd5703f8
And there will be no way to force converting a v2 cache back to v1, such
attempt will only clear free space tree and fallback to no space cache.
# mkfs.btrfs -f -O fst,^bgt test.img
# mount -o clear_cache,space_cache=v1 test.img /mnt/btrfs
# dmesg -t | tail -n 11
BTRFS warning: v1 space cache is deprecated, fallback to no space cache
BTRFS: device fsid f59daad2-3ab5-4f33-b752-a36cfb09b674 devid 1 transid 8 /dev/loop0 (7:0) scanned by mount (1419)
BTRFS info (device loop0): first mount of filesystem f59daad2-3ab5-4f33-b752-a36cfb09b674
BTRFS info (device loop0): using crc32c checksum algorithm
BTRFS info (device loop0): rebuilding free space tree
BTRFS info (device loop0): disabling free space tree
BTRFS info (device loop0): clearing compat-ro feature flag for FREE_SPACE_TREE (0x1)
BTRFS info (device loop0): clearing compat-ro feature flag for FREE_SPACE_TREE_VALID (0x2)
BTRFS info (device loop0): checking UUID tree
BTRFS info (device loop0): turning on async discard
BTRFS info (device loop0): force clearing of disk cache
# mount | grep /mnt/btrfs
/mnt/test.img on /mnt/btrfs type btrfs (rw,relatime,discard=async,nospace_cache,subvolid=5,subvol=/)
Signed-off-by: Qu Wenruo <wqu@suse.com>
Signed-off-by: David Sterba <dsterba@suse.com>
|
|
A swap file on btrfs will pin down block groups that cover the swap file
extent.
Pinned down block groups will be skipped for scrub and relocation.
These degradation on critical btrfs maintenance operations is never
properly educated to end users, and have already caused problems
including:
- Scrub finished too quick
Because the enabled swap file has pinned down most of the block
groups. Thus any file extents in those block groups, even not utilized
by the swap file, will be skipped from scrub.
- Unbalanced data and metadata usage, meanwhile relocation won't help
The same reason, pinned down block groups will not be considered as
relocation target, thus data extents that are not utilized by the swap
file can still be skipped from relocation.
Although we already have kernel messages for both scrub and balance, the
balance one is still info level.
To better communicate those potential long term problems, add the
following output into dmesg:
- Change the message level to warn for __btrfs_balance()
- Total pinned down block group number and size during swapfile activation
- Total released block group number and size during swapfile deactivation
The above messages have info level.
- The fact that pinned down block groups will not be scrubbed nor
balanced
The above message has warning level.
The example output would look like the following, for enabling a 1.2G
swapfile, which pinned down 2G block groups:
BTRFS info (device dm-3): swapfile activated on root 5 ino 257, pinned down 2147483648 bytes from 2 block group(s)
BTRFS warning (device dm-3): block groups with swapfile extents will not be scrubbed or balanced
Adding 1257468k swap on /mnt/btrfs/foobar. Priority:-1 extents:1 across:1257468k
BTRFS info (device dm-3): swapfile deactivated on root 5 ino 257, released 2147483648 bytes from 2 block group(s)
Reviewed-by: Filipe Manana <fdmanana@suse.com>
Signed-off-by: Qu Wenruo <wqu@suse.com>
Signed-off-by: David Sterba <dsterba@suse.com>
|
|
In Meta production, we have observed a large number of hosts running
kernels newer than 6.13 which hit hung tasks on
btrfs_read_folio()->lock_extents_for_read(). Looking through the history
in this codepath reveals an interesting history.
in 6.12, we merged
commit ac325fc2aad5 ("btrfs: do not hold the extent lock for entire read")
which holds the extent lock very narrowly while looking up the
extent_map. However, this proved to introduce a serious race with DIO
writes which was fixed in 6.14 with
commit acc18e1c1d8c0 ("btrfs: fix stale page cache after race between readahead and direct IO write")
That latter fix subtly changed the extent unlock point from the pre-6.12
regime. In 6.11, each read endio unlocked the extent it finished
reading, but in 6.14, the extent is locked/unlocked as a unit around the
entire readahead loop, while the individual folios are still unlocked as
the endios finish. This is mostly the same behavior, as all successful
reads will populate the page cache, so subsequent reads won't enter
btrfs and hit the extent lock. But in the case where the readahead
fails, perhaps because of a memory allocation failure doing compressed
reads, the page will not be brought up to date and a later read of an
overlapping range *will* block on the extent lock.
Why is this a problem?
On sufficiently large loaded systems, I have observed that direct
reclaim can run for minutes. Given that, consider two tasks on such a
system reading an overlapping range of a compressed file:
Task 1 locks the whole range and starts to read. Some allocation for
the compressed read for folio F fails and we carry on while holding the
extent lock for the full range.
Task 2 wants to read F, which is not uptodate and in page cache, so it
blocks on the extent lock held by Task 1.
Task 1 keeps getting stuck in direct reclaim (likely, we already
supposed an allocation failure above)
Task 2 stays blocked on the extent lock the whole time.
If you consider the effects of readahead_expand and imagine a file with
a 128k compressed extent followed by many smaller compressed extents,
you can imagine that the expanded window will result in subsequent reads
hitting many extents (128k/4k = 32) per lock window in the worst case.
The system likeley wouldn't be all that healthy anyway, so this is
likely not a critical improvement, but it does alleviate this one source
of stress and one thread's slowdown escalating to others.
To bring this behavior back to the old model, we should unlock the
extent at each loop of the readahead loop rather than in one shot at the
end. This allows such overlapping reads to proceed as they should.
Writes are fine because either the page has already been read and has an
appropriate state in the page cache to be invalidated (or not uptodate)
or it is still-to-be-read and the extent lock is still held protecting
it.
Reviewed-by: Filipe Manana <fdmanana@suse.com>
Reviewed-by: Qu Wenruo <wqu@suse.com>
Signed-off-by: Boris Burkov <boris@bur.io>
Signed-off-by: David Sterba <dsterba@suse.com>
|
|
The loop intends to copy the data in chunks up to 1M but we allocate the
pages array for the entire length and don't cap it to 1M. Fix this by
computing 'nr_pages' using 'copy_len' instead of 'length'.
While at it, also make 'nr_pages' and 'copy_len' const, as they never
change, to make the code more clear.
Reviewed-by: Johannes Thumshirn <johannes.thumshirn@wdc.com>
Signed-off-by: Filipe Manana <fdmanana@suse.com>
Signed-off-by: David Sterba <dsterba@suse.com>
|
|
When we set a root's reloc_root to NULL, we do it like this:
static void clear_reloc_root(struct btrfs_root *root)
{
root->reloc_root = NULL;
/*
* Need barrier to ensure clear_bit() only happens after
* root->reloc_root = NULL. Pairs with have_reloc_root().
*/
smp_wmb();
clear_bit(BTRFS_ROOT_DEAD_RELOC_TREE, &root->state);
}
So that a NULL reloc_root is always seen before seeing that the bit
BTRFS_ROOT_DEAD_RELOC_TREE was cleared.
But on the read side we have:
static bool reloc_root_is_dead(const struct btrfs_root *root)
{
smp_rmb();
if (test_bit(BTRFS_ROOT_DEAD_RELOC_TREE, &root->state))
return true;
return false;
}
And then callers of reloc_root_is_dead() access root->reloc_root.
Because the read memory barrier is placed before testing the bit, the CPU
is completely free to speculatively reorder those two loads. It can read
root->reloc_root before it actually checks the dead tree bit.
Sashiko reported this as an existing problem in another patch review, see
the link in the Link tag below.
Fix this by moving the read memory barrier to happen after testing the bit
and update the comment to reflect current reality.
Link: https://sashiko.dev/#/patchset/cf84f1a217c719e25b6b69e4298dd7afd36c9427.1781194426.git.fdmanana%40suse.com
Reviewed-by: Boris Burkov <boris@bur.io>
Signed-off-by: Filipe Manana <fdmanana@suse.com>
Signed-off-by: David Sterba <dsterba@suse.com>
|
|
[TEST FAILURE]
The test case generic/628 will fail if MOUNT_OPTIONS is set to
"-o nodatasum":
FSTYP -- btrfs
PLATFORM -- Linux/x86_64 btrfs-vm 7.1.0-rc4-custom+ #383 SMP PREEMPT_DYNAMIC Sat May 30 07:35:42 ACST 2026
MKFS_OPTIONS -- -O bgt -K /dev/mapper/test-scratch1
MOUNT_OPTIONS -- -o nodatasum /dev/mapper/test-scratch1 /mnt/scratch
generic/628 1s ... - output mismatch (see /home/adam/xfstests/results//generic/628.out.bad)
--- tests/generic/628.out 2022-05-11 11:25:30.816666664 +0930
+++ /home/adam/xfstests/results//generic/628.out.bad 2026-06-08 18:56:49.878542927 +0930
@@ -8,8 +8,9 @@
310f146ce52077fcd3308dcbe7632bb2 SCRATCH_MNT/a
310f146ce52077fcd3308dcbe7632bb2 SCRATCH_MNT/d
test reflink flag not set iflag
+XFS_IOC_CLONE: Invalid argument
310f146ce52077fcd3308dcbe7632bb2 SCRATCH_MNT/a
-310f146ce52077fcd3308dcbe7632bb2 SCRATCH_MNT/b
+d41d8cd98f00b204e9800998ecf8427e SCRATCH_MNT/b
...
[CAUSE]
The direct cause is that after "chattr +S", the btrfs inode will lose its
NODATASUM flag inherited from the mount option. E.g.:
# mkfs.btrfs -f $dev
# mount $dev $mnt -o nodatasum
# touch $mnt/foobar
# sync
# btrfs ins dump-tree -t 5 $dev | grep "(257 INODE_ITEM 0) itemoff" -A 3
item 4 key (257 INODE_ITEM 0) itemoff 15879 itemsize 160
generation 9 transid 9 size 0 nbytes 0
block group 0 mode 100644 links 1 uid 0 gid 0 rdev 0
sequence 1 flags 0x1(NODATASUM)
^^^^^^^^^ Proper NODATASUM flag
# chattr +S $mnt/foobar
# sync
# btrfs ins dump-tree -t 5 $dev | grep "(257 INODE_ITEM 0) itemoff" -A 3
item 4 key (257 INODE_ITEM 0) itemoff 15879 itemsize 160
generation 9 transid 10 size 0 nbytes 0
block group 0 mode 100644 links 1 uid 0 gid 0 rdev 0
sequence 2 flags 0x20(SYNC)
^^^^ Only the new SYNC flag
This makes the inode drop the old NODATASUM flag, while the new reflink
destination will still inherit the NODATASUM flag. The mismatching
NODATASUM flags will cause the reflink to fail.
The root cause is that, inside btrfs_fileattr_set() if no FS_NOCOW_FL is
set, we remove both NODATASUM and NODATACOW flag.
However we should not touch NODATASUM flag, as data COW doesn't require
checksum. Only NODATACOW implies NODATASUM, but DATACOW doesn't imply
DATASUM.
The deeper problems are:
- Fileattr API is too binary
It either clears or sets a flag, there is no "do not change" option.
So that why "chattr +S" implies "chattr -C", and is forcing us to
change NODATACOW along with NODATASUM flag.
- No way to change NODATASUM through fileattr API
In fact NODATASUM can only be modified through mount option.
The deeper problems are much harder to attack.
[FIX]
Remove NODATACOW flag when FS_NOCOW_FL is not set, but only remove
NODATASUM if "nodatasum" mount option is not set.
This allows the existing "chattr +C" then "chattr -C" to remove
both NODATACOW and NODATASUM flags on a default mount.
But for a mount with "nodatasum" option, the NODATASUM inode flag will
persist through either "chattr +C" and "chattr -C".
Fixes: 7e97b8daf634 ("btrfs: allow setting NOCOW for a zero sized file via ioctl")
Cc: stable@vger.kernel.org
Reviewed-by: Filipe Manana <fdmanana@suse.com>
Signed-off-by: Qu Wenruo <wqu@suse.com>
Reviewed-by: David Sterba <dsterba@suse.com>
Signed-off-by: David Sterba <dsterba@suse.com>
|
|
We are using 2 units for properties but we only set one property.
Fix this by using the correct amount: 1 unit.
Reviewed-by: Qu Wenruo <wqu@suse.com>
Signed-off-by: Filipe Manana <fdmanana@suse.com>
Reviewed-by: David Sterba <dsterba@suse.com>
Signed-off-by: David Sterba <dsterba@suse.com>
|
|
The qgroup ioctls update the quota tree, but they currently start their
transactions using the root of the inode passed to the ioctl. This makes
the transaction reservation depend on the path used for the ioctl instead
of the tree being modified.
Start qgroup ioctl transactions on the quota root instead. Take a reference
to fs_info->quota_root under qgroup_ioctl_lock before starting the
transaction, because quota disable can clear and put fs_info->quota_root
after the early quota-enabled check. Keep the reference until the
transaction handle is ended.
Suggested-by: Qu Wenruo <wqu@suse.com>
Reviewed-by: Qu Wenruo <wqu@suse.com>
Signed-off-by: Dongjiang Zhu <zhudongjiang@fnnas.com>
Reviewed-by: David Sterba <dsterba@suse.com>
Signed-off-by: David Sterba <dsterba@suse.com>
|
|
That function has the following problems:
- Read/write handling scattered across different locations
E.g. At the beginning there is a dedicated hole read handling, but
later short read handling is at an if() branch.
- Modifying of @pos and @length parameter for short read
Although it's completely fine to modify those parameters as they are
passed by value, but it can still be confusing to read. As normally
we would assume @pos and @length to be the original range.
But for short IO handling we modify @pos/@length, and completely
ignore @written.
- Unnecessary split for ordered extent and changeset handling
Both OE and changeset are only for writes, but they are handled in two
different if (write) {} blocks.
Refactor the function so that:
- Handling of reads and writes are concentrated in their code block
Now the handling of reads are in its own small if () branch.
Leaving the more complex writes handling to take the remaining
function, and reduce the indent level.
This also removes all unnecessary "if (write)" checks.
- Do not modify @pos and @length
Let short IO handling to manually calculate the remaining range.
Reviewed-by: Filipe Manana <fdmanana@suse.com>
Signed-off-by: Qu Wenruo <wqu@suse.com>
Signed-off-by: David Sterba <dsterba@suse.com>
|
|
This member records how many bytes are submitted for a direct
read/write, utilized by iomap_end() callback to handle short IO cases.
However iomap_end() callback is already providing an internally tracked
@written member, which is doing the same accounting and providing the
same value as btrfs_dio_data::submitted.
There is no need to duplicate the work, just remove btrfs_dio_data::submitted.
Reviewed-by: Filipe Manana <fdmanana@suse.com>
Signed-off-by: Qu Wenruo <wqu@suse.com>
Reviewed-by: David Sterba <dsterba@suse.com>
Signed-off-by: David Sterba <dsterba@suse.com>
|
|
sd_done() sense gate"
Yang Xiuwei <yangxiuwei@kylinos.cn> says:
This series fixes three resource-handling bugs in drivers/scsi/sd.c:
sd_probe() error cleanup, special_vec mempool leak on prep failure,
and sd_done() sense handling.
v1: https://lore.kernel.org/all/20260623100159.4018066-1-yangxiuwei@kylinos.cn/
Link: https://patch.msgid.link/20260707030333.22245-1-yangxiuwei@kylinos.cn
Signed-off-by: Martin K. Petersen (Oracle) <mkp@kernel.org>
|
|
Only enter the sense_key switch when the command returned CHECK
CONDITION with valid, non-deferred sense. The old condition let deferred
or invalid sense fall through and mis-handle the I/O.
Fixes: 03aba2f79594 ("[SCSI] sd/scsi_lib simplify sd_rw_intr and scsi_io_completion")
Reviewed-by: Damien Le Moal <dlemoal@kernel.org>
Signed-off-by: Yang Xiuwei <yangxiuwei@kylinos.cn>
Reviewed-by: Bart Van Assche <bvanassche@acm.org>
Link: https://patch.msgid.link/20260707030333.22245-4-yangxiuwei@kylinos.cn
Signed-off-by: Martin K. Petersen (Oracle) <mkp@kernel.org>
|
|
sd_set_special_bvec() allocates a special payload page for UNMAP and
WRITE SAME commands. If scsi_alloc_sgtables() fails afterward in
sd_setup_unmap_cmnd() or sd_setup_write_same{10,16}_cmnd(), the SCSI
midlayer does not call uninit_command() because RQF_DONTPREP is not set
yet, leaking the page.
Call sd_uninit_command() on error, and clear RQF_SPECIAL_PAYLOAD after
freeing the page.
Fixes: 81d926e8b552 ("sd: split sd_setup_discard_cmnd")
Reviewed-by: Damien Le Moal <dlemoal@kernel.org>
Signed-off-by: Yang Xiuwei <yangxiuwei@kylinos.cn>
Reviewed-by: John Garry <john.g.garry@oracle.com>
Link: https://patch.msgid.link/20260707030333.22245-3-yangxiuwei@kylinos.cn
Signed-off-by: Martin K. Petersen (Oracle) <mkp@kernel.org>
|
|
After device_add(&sdkp->disk_dev) succeeds, sd_large_pool_create()
failure must unregister disk_dev and let scsi_disk_release() free
sdkp. Going through out_free_index kfree()s an already registered device
and leaks the sysfs entry.
Fixes: 7179e626b76e ("scsi: sd: Enable sector size > PAGE_SIZE in SCSI sd driver")
Reviewed-by: Damien Le Moal <dlemoal@kernel.org>
Signed-off-by: Yang Xiuwei <yangxiuwei@kylinos.cn>
Reviewed-by: John Garry <john.g.garry@oracle.com>
Link: https://patch.msgid.link/20260707030333.22245-2-yangxiuwei@kylinos.cn
Signed-off-by: Martin K. Petersen (Oracle) <mkp@kernel.org>
|
|
After applying the commit associated with the Fixes tag, metadata file
buffers can be evicted from memory even after being marked dirty.
Consequently, operations such as rolling back sufile changes upon
error - which modify the buffer and were previously assumed incapable
of failure - can now fail.
This behavior causes syzbot to trigger a WARN_ON check immediately
following sufile function calls within the log writer.
Resolve this issue by introducing a macro, nilfs_sufile_warn_on_error(),
which uses WARN_ONCE to report unexpected errors only when the filesystem
has not degraded to read-only mode, returning -EIO or -EROFS accordingly.
Replace existing WARN_ON checks for unexpected errors following sufile
operations with this new macro.
Additionally, for nilfs_segctor_truncate_segments() - where an error must
be propagated to halt log writing if a sufile operation fails - modify
the function to return the error code appropriately.
Reported-by: syzbot+5957361606d7b750b874@syzkaller.appspotmail.com
Closes: https://syzkaller.appspot.com/bug?extid=5957361606d7b750b874
Fixes: 8c26c4e2694a ("nilfs2: fix issue with flush kernel thread after remount in RO mode because of driver's internal error or metadata corruption")
Cc: stable+noautosel@kernel.org # Warning suppression primarily; will request backport individually if needed
Signed-off-by: Ryusuke Konishi <konishi.ryusuke@gmail.com>
Signed-off-by: Viacheslav Dubeyko <slava@dubeyko.com>
|
|
synthesis
Modify bounds and union member access in mmap2 build_id synthesis. Bound
max_filename_len against the minimum of filename array capacity and the
outer union stack layout minus sample ID trailers. This prevents both
-E2BIG overruns and _FORTIFY_SOURCE array bounds aborts on strlcpy even
if the enclosing union expands.
Assisted-by: Antigravity:gemini-3.5-flash
Signed-off-by: Ian Rogers <irogers@google.com>
Signed-off-by: Namhyung Kim <namhyung@kernel.org>
|
|
module synthesis
Clamp long DSO names to mmap/mmap2 filename boundaries accounting for
sample ID headers to prevent buffer overruns in
perf_event__synthesize_modules_maps_cb(). Explicitly clear misc flags and
union padding to prevent stale Build-ID state from leaking between module
synthesis events, and cast event buffer pointers to avoid _FORTIFY_SOURCE
array bounds aborts when zeroing padding trailers.
Assisted-by: Antigravity:gemini-3.5-flash
Signed-off-by: Ian Rogers <irogers@google.com>
Signed-off-by: Namhyung Kim <namhyung@kernel.org>
|
|
Fix a pre-existing stack buffer overflow bug in
perf_event__synthesize_cgroup() where an in-place null padding loop wrote
bytes past the end of the cgrp_root stack array buffer during cgroup tree
traversal. Eliminate in-place path mutation, use PERF_ALIGN for path_len,
clamp raw_path_len to prevent sample ID header trailer overruns, and use
strlcpy with combined zero padding for alignment and sample ID headers.
Assisted-by: Antigravity:gemini-3.5-flash
Signed-off-by: Ian Rogers <irogers@google.com>
Signed-off-by: Namhyung Kim <namhyung@kernel.org>
|
|
in proc maps reader
Fix critical logic and boundary bugs in read_proc_maps_line() and caller.
Ensure any mid-line hex/dec/char parsing failure invokes io__drain_line()
safely, using a do-while loop to read and discard remaining characters
until a newline or EOF is reached. Clamp pathname extraction size to
account for trailing sample ID headers, use standard '//toolong' fallback
literal for over-length pathnames, emit timeout flags for truncated entries
securely via goto out;, and cast event buffer pointers to avoid
_FORTIFY_SOURCE array bounds aborts across synthesis handlers.
Assisted-by: Antigravity:gemini-3.5-flash
Signed-off-by: Ian Rogers <irogers@google.com>
Signed-off-by: Namhyung Kim <namhyung@kernel.org>
|
|
Use getline() to dynamically allocate the required line buffer for maps
parsing, guaranteeing bounds safety and avoiding compiler warnings
by evaluating the return value in the loop condition directly.
Assisted-by: Antigravity:gemini-3.5-flash
Signed-off-by: Ian Rogers <irogers@google.com>
Signed-off-by: Namhyung Kim <namhyung@kernel.org>
|
|
skip() ignores do_read()'s return value and unconditionally
subtracts the requested chunk size from 'size' on every iteration.
This was previously bounded by size being 'int': a maliciously
large 64-bit value was truncated on assignment, capping the loop
early by accident.
Now that size is size_t, a crafted file supplying a very large
size causes skip() to keep requesting BUFSIZ-sized reads and
subtracting BUFSIZ from size regardless of whether do_read()
actually succeeds, spinning indefinitely even after EOF or a read
error.
Check do_read()'s return value and break out of the loop on
failure or EOF, so forward progress is only counted when a read
actually succeeds.
Signed-off-by: Tanushree Shah <tshah@linux.ibm.com>
Signed-off-by: Namhyung Kim <namhyung@kernel.org>
|
|
read_ftrace_printk()/read_saved_cmdline()
Both functions read an attacker-controlled size directly from the
input file and pass size + 1 to malloc() before reading size bytes
into the result:
read_ftrace_printk(): size is an unsigned int from read4(). When
size == UINT_MAX, size + 1 overflows to 0, so malloc(0) returns a
minimal allocation while size itself remains UINT_MAX.
read_saved_cmdline(): size is an unsigned long long from read8().
When size == ULLONG_MAX, size + 1 overflows to 0 the same way.
In both cases, do_read(buf, size) then attempts to read the full,
unwrapped size into the tiny allocated buffer, a heap buffer
overflow.
This was previously masked by do_read()'s size parameter being
'int': passing these values truncated them, which the read()
syscall's own boundary checks rejected before any data was read.
Fixing that truncation (widening do_read() to size_t) is correct
on its own, but it removes this accidental protection and exposes
the pre-existing missing bounds check in both functions.
Reject the one value that causes the overflow before it's used, in
each function.
Signed-off-by: Tanushree Shah <tshah@linux.ibm.com>
Signed-off-by: Namhyung Kim <namhyung@kernel.org>
|
|
trace_event__cleanup()/trace_event__init()
trace_event__cleanup() frees t->pevent but never clears the
pointer. It can be called twice on the same trace_event: once
from trace_report()'s error path, and again from
perf_session__delete() during session teardown, resulting in a
double free / use-after-free.
Separately, trace_event__init() overwrites t->pevent/t->plugin_list
without releasing any existing handle, leaking memory if it's
called more than once on the same struct. eg. via a perf.data
file with multiple PERF_RECORD_HEADER_TRACING_DATA headers.
Guard against re-entry by returning early if t->pevent is already
NULL, and clear it after cleanup so a repeat call is a safe no-op.
Call trace_event__cleanup() at the start of trace_event__init(),
so a repeated init releases any existing handle before allocating
a new one.
Signed-off-by: Tanushree Shah <tshah@linux.ibm.com>
Signed-off-by: Namhyung Kim <namhyung@kernel.org>
|
|
The do_read() and skip() functions use 'int' for size parameters,
truncating 64-bit sizes from callers. This causes two issues:
1. Uninitialized memory dump: do_read() reads fewer bytes than
allocated, leaving uninitialized heap memory that gets written
to output files.
2. Out-of-bounds read: Parsing functions process the full 64-bit
size while only partial data was read into the buffer.
Change do_read(), __do_read(), and skip() to use size_t for size
parameters and ssize_t for return values (where applicable), matching
read()/write() system calls.
Update callers to use ssize_t for storing return values.
Fixes: 4a31e56599d4 ("perf tools: Get rid of read_or_die() in trace-event-read.c")
Signed-off-by: Tanushree Shah <tshah@linux.ibm.com>
Signed-off-by: Namhyung Kim <namhyung@kernel.org>
|
|
The USB receive path trusts the block header to contain the required
number of bytes and passes it to the reassembly routine. The routine
also trusts a malformed HCI packet type and can append more data than
the skb allocated from the advertised packet length. A malformed USB
transfer can therefore cause out-of-bounds reads or an skb tail
overwrite.
Validate block header availability, declared block size, packet type,
and reassembly tailroom. Drop the partial frame on an invalid block.
Signed-off-by: Li Qiang <liqiang01@kylinos.cn>
Signed-off-by: Luiz Augusto von Dentz <luiz.von.dentz@intel.com>
|
|
The reset path had two concurrency holes. Both are reachable in
practice when btintel_pcie_hw_error() is invoked from the HCI rx
path while another reset is being requested or is already in
flight.
1. data->reset_type was a plain shared field. The hw_error path
wrote it BEFORE the test_and_set_bit(RECOVERY_IN_PROGRESS)
guard inside btintel_pcie_reset(), so a second hw_error could
clobber the type chosen by an earlier in-flight request:
CPU0 (reset_work) CPU1 (hw_error #2)
dev_data->reset_type = PLDR
T2: read reset_type
dev_data->reset_type = FLR
reset() test_and_set sees 1
-> drops, but type already
clobbered
The hdev->reset callback (.reset = btintel_pcie_reset,
invoked via the sysfs reset attribute
/sys/class/bluetooth/hciX/reset and from hci_cmd_timeout())
compounded this by not writing reset_type at all -- it
inherited whatever value a previous hw_error / resume() had
left, which could be PLDR.
2. btintel_pcie_dump_debug_registers() was called unconditionally
at the top of hw_error(). When reset_work was already running
pci_try_reset_function(), the BT MMIO window can read all-1s
or trigger AER for the duration of the FLR, polluting the
debug dump with no useful information.
Refactor the reset path to make RECOVERY_IN_PROGRESS the sole
serializer for both the type write and the work scheduling:
- Replace btintel_pcie_reset(hdev) with
btintel_pcie_request_reset(data, type). The helper takes the
desired reset variant as a parameter and writes
data->reset_type only after winning test_and_set_bit(); losers
return without touching the field, so concurrent triggers can
no longer clobber an in-flight reset's type. reset_work()'s
read of reset_type is now ordered after the bit transition via
schedule_work()'s memory barrier.
- Add a thin btintel_pcie_hci_reset() wrapper for the
hdev->reset callback (invoked via the sysfs reset attribute
/sys/class/bluetooth/hciX/reset and from hci_cmd_timeout())
that always requests FLR explicitly, so these paths no longer
inherit stale state from prior error events.
- Add an early test_bit(RECOVERY_IN_PROGRESS) gate at the top of
hw_error() so dump_debug_registers() and the recovery-counter
bookkeeping are skipped when a reset is already in flight; the
authoritative test_and_set lives in request_reset() and races
cleanly against any caller that passes the optimistic check.
- Convert the two resume() reset sites (FREEZE/HIBERNATE and the
D0-error path) to request_reset(data, FLR), removing the
redundant manual reset_type writes.
Assisted-by: GitHub-Copilot:claude-4.7-opus
Signed-off-by: Kiran K <kiran.k@intel.com>
Signed-off-by: Luiz Augusto von Dentz <luiz.von.dentz@intel.com>
|
|
Do not export both functions since they are only used internally
within the bluetooth module.
Signed-off-by: Zijun Hu <zijun.hu@oss.qualcomm.com>
Signed-off-by: Luiz Augusto von Dentz <luiz.von.dentz@intel.com>
|
|
Wiko Hi MateBook 14 Ryzen 200 laptops (DMI system-product-name
"MNCA-XX", board "M1060") are equipped with an RTL8852BE Wi-Fi/BT
combo chip (rtw89_8852be), whose Bluetooth radio enumerates as
1357:c123 instead of one of the already-supported 1358:c123 / 0bda:c123
identifiers, presumably due to OEM rebranding. Without a matching
entry it only matches the generic USB Bluetooth class fallback, so the
Realtek firmware/config (rtl8852btu_fw.bin / rtl8852btu_config.bin) is
never loaded and the adapter cannot discover or connect to any device,
even though hciconfig reports it as powered and scanning.
Device descriptor:
idVendor 0x1357
idProduct 0xc123
bcdDevice 0.00
iManufacturer 1 Realtek
iProduct 2 Bluetooth Radio
bDeviceClass 224 Wireless
bDeviceSubClass 1 Radio Frequency
bDeviceProtocol 1 Bluetooth
Adding the same BTUSB_REALTEK | BTUSB_WIDEBAND_SPEECH quirk already
used for 1358:c123 and 0bda:c123 fixes firmware loading and normal
operation.
Signed-off-by: Pavel Zverev <playximik29@gmail.com>
Signed-off-by: Luiz Augusto von Dentz <luiz.von.dentz@intel.com>
|
|
btintel_pcie_coredump_worker() handled three unrelated jobs in one
work item: collect a DRAM trace coredump, read the hardware exception
event, and read the firmware-trigger event. The worker walked three
flag bits at runtime and each interrupt path mutated multiple bits
to communicate which sub-jobs the worker should run, which made the
ownership rules for those bits hard to reason about and entangled
the trigger reason with the in-progress accounting.
Replace the single combined worker with three single-purpose ones,
each owning exactly one flag:
coredump_work -> btintel_pcie_dump_traces()
guarded by COREDUMP_INPROGRESS
hwexp_work -> btintel_pcie_read_hwexp()
guarded by CORE_HALTED (already permanent until
re-probe; HWEXP_INPROGRESS is now redundant
and removed)
fwtrigger_work -> btintel_pcie_dump_fwtrigger_event()
guarded by FWTRIGGER_DUMP_INPROGRESS
All three workers are queued on a shared ordered workqueue (renamed
coredump_workqueue -> dump_workqueue) so a companion event reader
(hwexp/fwtrigger) and the coredump always run FIFO. Companion work
is queued before coredump_work so dmp_hdr.event_type/event_id are
populated by the time dump_traces() consumes them, preserving the
original ordering.
Introduce btintel_pcie_queue_coredump() to centralize the coredump
trigger contract: it is the single writer of COREDUMP_INPROGRESS and
of dmp_hdr.trigger_reason, sets both atomically against concurrent
triggers, and rolls back the bit if the workqueue is disabled
(reset/remove in progress) so a later trigger after re-probe can
succeed. All four trigger sites (HWEXP IRQ, FW-trigger IRQ,
devcoredump user trigger, resume() D0 error path) go through the
helper.
Per-work guard bits are now cleared at the tail of each worker
rather than in the middle of the combined worker, which closes a
subtle race where a duplicate IRQ could observe a cleared bit and
requeue while the previous pass was still finalizing
dev_coredumpv().
reset_work() and remove() now disable_work_sync() all three workers
and, on the FLR-failure path, enable_work() all three to keep their
disable counters balanced. The PLDR/FLR-success contract (re-probe
re-INIT_WORKs everything with counter 0) is preserved.
No functional change to the dump payloads; this is a pure
restructuring of the worker dispatch and its synchronization.
Signed-off-by: Kiran K <kiran.k@intel.com>
Assisted-by: GitHub-Copilot:claude-4.7-opus
Signed-off-by: Luiz Augusto von Dentz <luiz.von.dentz@intel.com>
|
|
In nxp_serdev_probe(), if hci_register_dev() succeeds but ps_setup()
fails, the error path jumps to 'probe_fail' which only calls
hci_free_dev() and asserts the reset GPIO, but does NOT call
hci_unregister_dev() first.
This leaves the HCI device registered in the system with its backing
memory freed, leading to a use-after-free when userspace subsequently
accesses the device (e.g. via hciconfig or bluetoothd).
Fix by adding a 'probe_fail_unregister' label that calls
hci_unregister_dev() before falling through to the existing
'probe_fail' label. The original 'probe_fail' label is preserved
for the case where hci_register_dev() itself fails (device was
never registered, so no unregister is needed).
Signed-off-by: Zhao Dongdong <zhaodongdong@kylinos.cn>
Reviewed-by: Neeraj Sanjay Kale <neeraj.sanjaykale@nxp.com>
Signed-off-by: Luiz Augusto von Dentz <luiz.von.dentz@intel.com>
|
|
Add the vendor/product ID (0x0b05, 0x1d70) to the usb_device_id table for
the Realtek RTL8761CU-based ASUS USB-BT600 adapter. It binds via the
generic Bluetooth class today, so BTUSB_REALTEK is never set and the
rtl8761cu firmware is not loaded, leaving the controller non-functional.
With the entry the driver loads rtl_bt/rtl8761cu_fw.bin (already shipped by
linux-firmware) and the adapter works (tested: A2DP and ASHA).
Similar to commit bc597f0cc44f
("Bluetooth: btusb: Add TP-Link UB600 for Realtek 8761BUV").
Device info from /sys/kernel/debug/usb/devices:
T: Bus=01 Lev=01 Prnt=01 Port=01 Cnt=01 Dev#= 23 Spd=12 MxCh= 0
D: Ver= 1.10 Cls=e0(wlcon) Sub=01 Prot=01 MxPS=64 #Cfgs= 1
P: Vendor=0b05 ProdID=1d70 Rev= 2.00
S: Manufacturer=Realtek
S: Product=Bluetooth Controller
C:* #Ifs= 2 Cfg#= 1 Atr=e0 MxPwr=100mA
I:* If#= 0 Alt= 0 #EPs= 3 Cls=e0(wlcon) Sub=01 Prot=01 Driver=btusb
E: Ad=81(I) Atr=03(Int.) MxPS= 64 Ivl=1ms
E: Ad=02(O) Atr=02(Bulk) MxPS= 64 Ivl=0ms
E: Ad=82(I) Atr=02(Bulk) MxPS= 64 Ivl=0ms
I:* If#= 1 Alt= 0 #EPs= 2 Cls=e0(wlcon) Sub=01 Prot=01 Driver=btusb
E: Ad=83(I) Atr=01(Isoc) MxPS= 0 Ivl=1ms
E: Ad=03(O) Atr=01(Isoc) MxPS= 0 Ivl=1ms
I: If#= 1 Alt= 1 #EPs= 2 Cls=e0(wlcon) Sub=01 Prot=01 Driver=btusb
E: Ad=83(I) Atr=01(Isoc) MxPS= 9 Ivl=1ms
E: Ad=03(O) Atr=01(Isoc) MxPS= 9 Ivl=1ms
I: If#= 1 Alt= 2 #EPs= 2 Cls=e0(wlcon) Sub=01 Prot=01 Driver=btusb
E: Ad=83(I) Atr=01(Isoc) MxPS= 17 Ivl=1ms
E: Ad=03(O) Atr=01(Isoc) MxPS= 17 Ivl=1ms
I: If#= 1 Alt= 3 #EPs= 2 Cls=e0(wlcon) Sub=01 Prot=01 Driver=btusb
E: Ad=83(I) Atr=01(Isoc) MxPS= 25 Ivl=1ms
E: Ad=03(O) Atr=01(Isoc) MxPS= 25 Ivl=1ms
I: If#= 1 Alt= 4 #EPs= 2 Cls=e0(wlcon) Sub=01 Prot=01 Driver=btusb
E: Ad=83(I) Atr=01(Isoc) MxPS= 33 Ivl=1ms
E: Ad=03(O) Atr=01(Isoc) MxPS= 33 Ivl=1ms
I: If#= 1 Alt= 5 #EPs= 2 Cls=e0(wlcon) Sub=01 Prot=01 Driver=btusb
E: Ad=83(I) Atr=01(Isoc) MxPS= 49 Ivl=1ms
E: Ad=03(O) Atr=01(Isoc) MxPS= 49 Ivl=1ms
I: If#= 1 Alt= 6 #EPs= 2 Cls=e0(wlcon) Sub=01 Prot=01 Driver=btusb
E: Ad=83(I) Atr=01(Isoc) MxPS= 63 Ivl=1ms
E: Ad=03(O) Atr=01(Isoc) MxPS= 63 Ivl=1ms
Cc: stable@vger.kernel.org
Signed-off-by: Christoph Zwerschke <cito@online.de>
Reviewed-by: Paul Menzel <pmenzel@molgen.mpg.de>
Signed-off-by: Luiz Augusto von Dentz <luiz.von.dentz@intel.com>
|
|
Add the vendor/product ID (0x0b05, 0x1bef) to the usb_device_id table for
the Realtek RTL8761CU-based ASUS USB-BT540 adapter. It binds via the
generic Bluetooth class today, so BTUSB_REALTEK is never set and the
rtl8761cu firmware is not loaded, leaving the controller non-functional.
With the entry the driver loads rtl_bt/rtl8761cu_fw.bin (already shipped by
linux-firmware) and the adapter works (tested: A2DP and ASHA).
Similar to commit bc597f0cc44f
("Bluetooth: btusb: Add TP-Link UB600 for Realtek 8761BUV").
Device info from /sys/kernel/debug/usb/devices:
T: Bus=01 Lev=01 Prnt=01 Port=01 Cnt=01 Dev#= 22 Spd=12 MxCh= 0
D: Ver= 1.10 Cls=e0(wlcon) Sub=01 Prot=01 MxPS=64 #Cfgs= 1
P: Vendor=0b05 ProdID=1bef Rev= 2.00
S: Manufacturer=Realtek
S: Product=Bluetooth Controller
C:* #Ifs= 2 Cfg#= 1 Atr=e0 MxPwr=100mA
I:* If#= 0 Alt= 0 #EPs= 3 Cls=e0(wlcon) Sub=01 Prot=01 Driver=btusb
E: Ad=81(I) Atr=03(Int.) MxPS= 64 Ivl=1ms
E: Ad=02(O) Atr=02(Bulk) MxPS= 64 Ivl=0ms
E: Ad=82(I) Atr=02(Bulk) MxPS= 64 Ivl=0ms
I:* If#= 1 Alt= 0 #EPs= 2 Cls=e0(wlcon) Sub=01 Prot=01 Driver=btusb
E: Ad=83(I) Atr=01(Isoc) MxPS= 0 Ivl=1ms
E: Ad=03(O) Atr=01(Isoc) MxPS= 0 Ivl=1ms
I: If#= 1 Alt= 1 #EPs= 2 Cls=e0(wlcon) Sub=01 Prot=01 Driver=btusb
E: Ad=83(I) Atr=01(Isoc) MxPS= 9 Ivl=1ms
E: Ad=03(O) Atr=01(Isoc) MxPS= 9 Ivl=1ms
I: If#= 1 Alt= 2 #EPs= 2 Cls=e0(wlcon) Sub=01 Prot=01 Driver=btusb
E: Ad=83(I) Atr=01(Isoc) MxPS= 17 Ivl=1ms
E: Ad=03(O) Atr=01(Isoc) MxPS= 17 Ivl=1ms
I: If#= 1 Alt= 3 #EPs= 2 Cls=e0(wlcon) Sub=01 Prot=01 Driver=btusb
E: Ad=83(I) Atr=01(Isoc) MxPS= 25 Ivl=1ms
E: Ad=03(O) Atr=01(Isoc) MxPS= 25 Ivl=1ms
I: If#= 1 Alt= 4 #EPs= 2 Cls=e0(wlcon) Sub=01 Prot=01 Driver=btusb
E: Ad=83(I) Atr=01(Isoc) MxPS= 33 Ivl=1ms
E: Ad=03(O) Atr=01(Isoc) MxPS= 33 Ivl=1ms
I: If#= 1 Alt= 5 #EPs= 2 Cls=e0(wlcon) Sub=01 Prot=01 Driver=btusb
E: Ad=83(I) Atr=01(Isoc) MxPS= 49 Ivl=1ms
E: Ad=03(O) Atr=01(Isoc) MxPS= 49 Ivl=1ms
I: If#= 1 Alt= 6 #EPs= 2 Cls=e0(wlcon) Sub=01 Prot=01 Driver=btusb
E: Ad=83(I) Atr=01(Isoc) MxPS= 63 Ivl=1ms
E: Ad=03(O) Atr=01(Isoc) MxPS= 63 Ivl=1ms
Cc: stable@vger.kernel.org
Signed-off-by: Christoph Zwerschke <cito@online.de>
Reviewed-by: Paul Menzel <pmenzel@molgen.mpg.de>
Signed-off-by: Luiz Augusto von Dentz <luiz.von.dentz@intel.com>
|
|
nokia_setup_fw() walks a length-prefixed firmware stream and
decodes HCI command packets from each record.
Check that each record fits in the remaining firmware image, that command
records contain the HCI command header, and that the payload length is
covered before submitting the command.
Signed-off-by: Pengpeng Hou <pengpeng@iscas.ac.cn>
Signed-off-by: Luiz Augusto von Dentz <luiz.von.dentz@intel.com>
|
|
Add the USB ID (13d3:3503) for the IMC Networks Qualcomm Atheros
QCA9377 Bluetooth controller to the btusb quirks table. This device
requires Qualcomm Rome firmware and wideband speech support to function
properly; otherwise, BLE scanning fails with HCI unexpected event
opcode 0x2005 errors.
The device reports the following in /sys/kernel/debug/usb/devices:
P: Vendor=13d3 ProdID=3503 Rev= 0.01
C:* #Ifs= 2 Cfg#= 1 Atr=e0 MxPwr=100mA
I:* If#= 0 Alt= 0 #EPs= 3 Cls=e0(wlcon) Sub=01 Prot=01 Driver=btusb
E: Ad=81(I) Atr=03(Int.) MxPS= 16 Ivl=1ms
E: Ad=82(I) Atr=02(Bulk) MxPS= 64 Ivl=0ms
E: Ad=02(O) Atr=02(Bulk) MxPS= 64 Ivl=0ms
I:* If#= 1 Alt= 0 #EPs= 2 Cls=e0(wlcon) Sub=01 Prot=01 Driver=btusb
E: Ad=83(I) Atr=01(Isoc) MxPS= 0 Ivl=1ms
E: Ad=03(O) Atr=01(Isoc) MxPS= 0 Ivl=1ms
I: If#= 1 Alt= 1 #EPs= 2 Cls=e0(wlcon) Sub=01 Prot=01 Driver=btusb
E: Ad=83(I) Atr=01(Isoc) MxPS= 9 Ivl=1ms
E: Ad=03(O) Atr=01(Isoc) MxPS= 9 Ivl=1ms
I: If#= 1 Alt= 2 #EPs= 2 Cls=e0(wlcon) Sub=01 Prot=01 Driver=btusb
E: Ad=83(I) Atr=01(Isoc) MxPS= 17 Ivl=1ms
E: Ad=03(O) Atr=01(Isoc) MxPS= 17 Ivl=1ms
I: If#= 1 Alt= 3 #EPs= 2 Cls=e0(wlcon) Sub=01 Prot=01 Driver=btusb
E: Ad=83(I) Atr=01(Isoc) MxPS= 25 Ivl=1ms
E: Ad=03(O) Atr=01(Isoc) MxPS= 25 Ivl=1ms
I: If#= 1 Alt= 4 #EPs= 2 Cls=e0(wlcon) Sub=01 Prot=01 Driver=btusb
E: Ad=83(I) Atr=01(Isoc) MxPS= 33 Ivl=1ms
E: Ad=03(O) Atr=01(Isoc) MxPS= 33 Ivl=1ms
I: If#= 1 Alt= 5 #EPs= 2 Cls=e0(wlcon) Sub=01 Prot=01 Driver=btusb
E: Ad=83(I) Atr=01(Isoc) MxPS= 49 Ivl=1ms
E: Ad=03(O) Atr=01(Isoc) MxPS= 49 Ivl=1ms
Signed-off-by: Tibor Harcsa <silurust@gmail.com>
Signed-off-by: Luiz Augusto von Dentz <luiz.von.dentz@intel.com>
|