| Age | Commit message (Collapse) | Author | Files | Lines |
|
SMB2 QUERY_INFO security requests currently return owner, group, and DACL
information without checking the access granted to the opened handle. A
handle opened with only SYNCHRONIZE or READ_ATTRIBUTES can consequently
read the security descriptor.
Require READ_CONTROL when OWNER_SECINFO, GROUP_SECINFO, or DACL_SECINFO
is requested and return STATUS_ACCESS_DENIED otherwise.
This fixes smb2.getinfo.getinfo_access.
Signed-off-by: Namjae Jeon <linkinjeon@kernel.org>
|
|
FILE_NORMALIZED_NAME_INFORMATION is not handled and is returned as
STATUS_INVALID_INFO_CLASS. SMB 3.1.1 clients use this information class
to obtain the share-relative path with the on-disk name casing.
Build the normalized path from the opened dentry, remove the leading
share-relative separator, and recover the canonical named-stream casing
from its backing xattr. Return an empty name for the share root and
STATUS_NOT_SUPPORTED for dialects older than SMB 3.1.1.
Also distinguish a named $DATA stream on a directory from the unnamed
data stream so that directory:stream:$DATA can be opened normally.
This fixes smb2.getinfo.normalized.
Signed-off-by: Namjae Jeon <linkinjeon@kernel.org>
|
|
SMB2 QUERY_INFO security requests with an output buffer too small for
the self-relative security descriptor header can fall through descriptor
construction and be reported as STATUS_INVALID_INFO_CLASS.
After validating the file handle, reject buffers shorter than struct
smb_ntsd with STATUS_BUFFER_TOO_SMALL before building the descriptor.
This fixes smb2.getinfo.qsec_buffercheck.
Signed-off-by: Namjae Jeon <linkinjeon@kernel.org>
|
|
Variable-length file information handlers use the client output length
while constructing the response. FILE_ALL_INFORMATION can consequently
return -EINVAL before the common buffer check, while stream information
can stop building the complete result too early.
Build the complete response within the available server response buffer
and apply the client output length only when selecting the final status
and transmitted length. Use the protocol-defined fixed sizes for all,
alternate-name, and stream information to distinguish
STATUS_INFO_LENGTH_MISMATCH from STATUS_BUFFER_OVERFLOW.
This fixes smb2.getinfo.qfile_buffercheck.
Signed-off-by: Namjae Jeon <linkinjeon@kernel.org>
|
|
check_lock_range() uses inclusive ranges. Its callers pass the end
offset as start + length - 1, so start == end represents a valid
single-byte range rather than an empty range.
The start == end shortcut therefore skips mandatory byte-range lock
checks for one-byte reads, writes, copychunk operations and one-byte
truncate ranges. A conflicting lock covering that byte is not checked
and the operation is allowed to proceed.
Remove the shortcut. The truncate size == inode->i_size case is already
handled by only calling check_lock_range() when the new size differs
from the current file size.
Fixes: 5d510ac31626 ("ksmbd: skip lock-range check on equal size to avoid size==0 underflow")
Signed-off-by: Guangshuo Li <lgs201920130244@gmail.com>
Signed-off-by: Namjae Jeon <linkinjeon@kernel.org>
|
|
The query-info buffer check returns STATUS_INFO_LENGTH_MISMATCH for
every output buffer smaller than the complete response. Variable-length
filesystem information instead requires STATUS_BUFFER_OVERFLOW when the
fixed portion fits but the complete data does not.
Pass the fixed size for each filesystem information class to the buffer
checker. Keep INFO_LENGTH_MISMATCH for buffers below that size, and
return BUFFER_OVERFLOW with a response truncated to the requested length
for larger partial buffers.
This fixes smb2.getinfo.qfs_buffercheck.
Signed-off-by: Namjae Jeon <linkinjeon@kernel.org>
|
|
Named streams are stored as extended attributes on the base inode. The
VFS read and write helpers reject directory inodes before or together
with checking whether the handle represents a stream.
Permit read and write operations when a directory-backed handle is a
named stream. Continue rejecting direct I/O on ordinary directory
handles.
This fixes creation of the directory stream in smb2.getinfo.complex.
Signed-off-by: Namjae Jeon <linkinjeon@kernel.org>
|
|
SMB clients can currently create an EA named NTACL because SMB EAs are
mapped into the user namespace while the ksmbd security descriptor is
stored as security.NTACL. Allowing the reserved logical name makes the
server-private ACL metadata appear writable through the SMB EA API.
Reject NTACL, DOSATTRIB, and DosStream-prefixed EA names without regard
to case. Filter the same private names from EA query results so stale or
externally-created user namespace attributes cannot be exposed.
This fixes smb2.ea.acl_xattr when acl_xattr_name is configured as
NTACL.
Signed-off-by: Namjae Jeon <linkinjeon@kernel.org>
|
|
DELETE_ON_CLOSE is currently accepted for files carrying the read-only
DOS attribute. The server consequently creates or opens the file and
marks it for deletion instead of returning STATUS_CANNOT_DELETE.
Reject creation of a new read-only file with DELETE_ON_CLOSE. For an
existing file, load the stored DOS attributes before accepting the
create option. Also reject FileDispositionInformation when the opened
file has the read-only attribute.
Preserve the explicit STATUS_CANNOT_DELETE value while unwinding the
CREATE request.
This fixes smb2.delete-on-close-perms.READONLY.
Signed-off-by: Namjae Jeon <linkinjeon@kernel.org>
|
|
The SMB2 create maximal-access context is currently calculated from
POSIX mode bits when the client does not request MAXIMUM_ALLOWED. This
overwrites the access granted by a stored Windows DACL.
Calculate the create-context result with the DACL permission checker.
Recognize the S-1-3-4 Owner Rights SID as applying to the object owner
and process its allow and deny ACEs in ACL order.
When an Owner Rights ACE is present, do not add the owner implicit
READ_CONTROL and WRITE_DAC rights. The Owner Rights ACE replaces those
implicit grants as required by Windows access-check semantics.
Without an Owner Rights ACE, preserve the existing implicit owner grants,
including FILE_READ_ATTRIBUTES and DELETE.
This fixes smb2.acls.OWNER-RIGHTS and its deny variants without regressing
smb2.acls.GENERIC.
Signed-off-by: Namjae Jeon <linkinjeon@kernel.org>
|
|
SMB shares can advertise access-based directory enumeration. ksmbd does
not currently provide a share option or filter inaccessible directory
entries.
Add a hide-unreadable share flag and advertise
SMB2_SHAREFLAG_ACCESS_BASED_DIRECTORY_ENUM when it is enabled. During
QUERY_DIRECTORY, omit entries unless the connected user has
FILE_READ_DATA, FILE_READ_EA, and FILE_READ_ATTRIBUTES access according
to the Windows ACL.
Keep the existing implicit access allowances for normal CREATE
permission checks while using strict access-mask matching for directory
enumeration.
Signed-off-by: Namjae Jeon <linkinjeon@kernel.org>
|
|
The DACL permission check looks for an ACE matching the current user and
falls back to the Everyone ACE. It does not consider an Authenticated
Users ACE, even though an authenticated session is a member of that
well-known group.
As a result, opening a file whose access is granted through S-1-5-11 can
incorrectly fail with STATUS_ACCESS_DENIED. Treat an Authenticated Users
ACE as a fallback entry alongside Everyone.
The maximal access calculation also combines access masks from every ACE,
regardless of whether its SID applies to the current user. This can grant
rights belonging to an unrelated principal. Process only ACEs applying to
the user, Everyone, or Authenticated Users, and accumulate allowed and
denied masks in ACL order. Preserve explicitly requested access bits so
they are validated against the resulting maximal mask.
When ACCESS_SYSTEM_SECURITY is denied, report STATUS_PRIVILEGE_NOT_HELD
instead of the generic STATUS_ACCESS_DENIED. Access to the system ACL
requires a security privilege that ksmbd does not grant.
For regular files, include FILE_EXECUTE in maximal access when the client
requested GENERIC_EXECUTE and the DACL grants the complete file-read set.
Keep a direct FILE_EXECUTE request subject to the explicit DACL bit. This
matches the POSIX file ACL mapping without broadening specific execute
requests.
Do not replace rights from an applicable NT ACE with a POSIX ACL entry.
The POSIX ACL is only a fallback when no user, Everyone, or Authenticated
Users ACE applies; otherwise it can incorrectly broaden the stored DACL.
This fixes smb2.maximum_allowed.maximum_allowed.
Signed-off-by: Namjae Jeon <linkinjeon@kernel.org>
|
|
An SMB2 WRITE request with a negative offset returns -EINVAL directly
from smb2_write(). This bypasses the common error response path, leaving
the client waiting until the request times out.
ksmbd also allows nonempty writes at or beyond MAXFILESIZE as defined by
[MS-FSA]. Writes beyond the limit must fail with
STATUS_INVALID_PARAMETER. Writes ending at the limit fail with
STATUS_DISK_FULL, while a zero-length write remains valid.
Route negative offsets through the common error path and validate the end
offset of nonempty writes against MAXFILESIZE.
This fixes smb2.rw.invalid.
Signed-off-by: Namjae Jeon <linkinjeon@kernel.org>
|
|
SMB3.1.1 multichannel connections belonging to the same session must use
the same negotiated encryption cipher.
ksmbd validates the dialect and client GUID during session binding, but
does not compare the cipher negotiated by the new connection with the
cipher used by the existing session channels. This allows a channel
negotiated with AES-128-CCM to bind to a session using AES-128-GCM.
Compare the new connection's cipher with an existing session channel and
return STATUS_INVALID_PARAMETER when they differ.
This fixes smb2.session.bind_negative_smb3encGtoCs.
Signed-off-by: Namjae Jeon <linkinjeon@kernel.org>
|
|
XFS already defers some writes to a workqueue when transactions are
needed to process the I/O completion. Disable the block layer bio task
completion in this case to avoid a major performance drop.
Fixes: efbde6f9f449 ("iomap: use BIO_COMPLETE_IN_TASK for dropbehind writeback")
Link: https://lore.kernel.org/all/8124341f-3af2-4a16-897d-38db5ab5a9d4@columbia.edu/
Signed-off-by: Tal Zussman <tz2294@columbia.edu>
Reviewed-by: "Darrick J. Wong" <djwong@kernel.org>
Reviewed-by: Christoph Hellwig <hch@lst.de>
Link: https://patch.msgid.link/20260810-xfs-dontcache-double-defer-v1-1-aea7484b3e49@columbia.edu
Signed-off-by: Jens Axboe <axboe@kernel.dk>
|
|
futex_exit_release() and futex_exec_release() are identical now. That means
also exit_mm_release() and exec_mm_release() are identical.
Consolidate the whole lot and remove the redundant copies.
Signed-off-by: Thomas Gleixner <tglx@kernel.org>
Reviewed-by: Kyle Zeng <kylebot@openai.com>
Acked-by: Peter Zijlstra <peterz@infradead.org>
|
|
The check for private futexes whether the waiter's mm, which is stored in
the futex_key and copied into the pi_state, is the same as the owner's mm
is not sufficient for exec(). exec() has a gap where the mm check fails to
give the correct answer:
exec()
...
exec_release_mm()
futex_exec_release()
tsk::futex::exit_state = EXITING;
cleanup_robust_list();
1) tsk::futex::exit_state = OK;
...
old_mm = tsk::mm;
2) tsk::mm = ->mm;
Between #1 and #2 the check for the mm is wrong as that mm is about to be
swapped out and eventually freed.
Plug this gap by:
1) Setting tsk::futex::exit_state to FUTEX_STATE_DEAD in
futex_exec_release()
2) Setting tsk::futex::exit_state to FUTEX_STATE_OK after
the mm has been switched.
From a futex point of view the task is dead after it finished the robust
list cleanup up to the point where it sets the state to OK again.
Fixes: 80367ad01d93 ("futex: Add basic infrastructure for local task local hash")
Signed-off-by: Thomas Gleixner <tglx@kernel.org>
Reviewed-by: Kyle Zeng <kylebot@openai.com>
Acked-by: Peter Zijlstra <peterz@infradead.org>
Cc: stable@vger.kernel.org
|
|
kernfs_iop_listxattr() only needs to report existing xattrs, but it uses
kernfs_iattrs(), which allocates kernfs_iattrs when the node does not
have one yet.
This makes a query operation create persistent per-node metadata even
when the xattr list is empty.
Use kernfs_iattrs_noalloc() instead and return an empty list when no
iattrs exist.
Signed-off-by: Yichong Chen <chenyichong@uniontech.com>
Acked-by: Tejun Heo <tj@kernel.org>
Link: https://patch.msgid.link/20260731120554.630147-1-chenyichong@uniontech.com
Signed-off-by: Danilo Krummrich <dakr@kernel.org>
|
|
Pull ceph fixes from Ilya Dryomov:
"A handful of tiny fixes, with the main ones being a follow-up for
CEPH_IOC_SET_LAYOUT{,_POLICY} ioctl permissions check that went into
rc5 and a userspace compatibility fixup. The rest mostly harden
against malformed network input. All marked for stable"
* tag 'ceph-for-7.2-rc8' of https://github.com/ceph/ceph-client:
ceph: use the mount idmap for the owner checks in the SET_LAYOUT ioctls
ceph: fix MDS random selection readiness predicate
libceph: Avoid using invalid osd indices from primary_temp
libceph: fix OOB read in decode_watchers() via missing bounds check
libceph: fix multiple unsafe decodes in decode_locker()
libceph: tolerate addrvecs with multiple entries of the same type
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/vfs/vfs
Pull vfs fixes from Christian Brauner:
- Don't warn when a mount is completed from another user namespace.
fsopen() records the caller's user namespace in fc->user_ns and
hands back an ordinary file descriptor. The task that calls
fsconfig(FSCONFIG_CMD_CREATE) doesn't have to be the one that
created the context, and mount_capable() lets it through as long
as the caller has CAP_SYS_ADMIN over fc->user_ns, which anyone in
an ancestor namespace does. So fc->user_ns != current_user_ns()
is something an unprivileged user can arrange.
Both overlayfs and binfmt_misc WARN_ON() that. Overlayfs already
has the same check as a plain error return in ovl_parse_param().
Drop the WARN_ON() and just refuse. Add selftests for both cases.
- Reject pid allocations through dead ancestor pid namespaces.
Require PIDNS_ADDING in every namespace that will receive the pid
before publishing any of them. That preserves the invariant that
free_pid() never decrements pid_allocated in a namespace whose
child_reaper is no longer live. The existing ENOMEM behavior is
unchanged.
* tag 'vfs-7.2-rc8.fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/vfs/vfs:
pid: reject allocations through dead ancestor pid namespaces
selftests/filesystems: test completing a context from another user namespace
binfmt_misc: don't warn when the mount is completed from another user namespace
ovl: don't warn when the mount is completed from another user namespace
|
|
CONFIG_NR_CPUS doesn't define on some UP platforms (e.g. arm), so this
can cause make oldconfig to loop indefinitely when CONFIG_SMP=n:
$ make ARCH=arm allmodconfig
$ sed -i "/CONFIG_SMP=y/d" .config
$ sed -i "/CONFIG_EROFS_FS_ZIP_LZMA_DEFAULT_MAX_STREAMS.*/d" .config
EROFS LZMA default maximum decompression streams (EROFS_FS_ZIP_LZMA_DEFAULT_MAX_STREAMS) [0] (NEW)
EROFS LZMA default maximum decompression streams (EROFS_FS_ZIP_LZMA_DEFAULT_MAX_STREAMS) [0] (NEW)
...
Let's guard NR_CPUS with SMP instead of using a hardcoded arbitrary CPU
uplimit here, similar to commit a3344078101c ("mm: make SPLIT_PTE_PTLOCKS
depend on SMP").
The initial report from SJ Park was for m68k [1] (m68k is the only arch
without NR_CPUS in Kconfig), and that got fixed in commit 1fd495ef09ee
("m68k: Define NR_CPUS to 1")
Reported-by: SJ Park <sj@kernel.org>
Link: https://lore.kernel.org/all/anuyFHLUGDjZWY4K@XiangdeMacBook-Pro.local/T/#u [1]
Closes: https://lore.kernel.org/r/20260728065447.91511-1-sj@kernel.org
Reported-by: Guenter Roeck <groeck7@gmail.com>
Closes: https://lore.kernel.org/r/87853c96-cc8f-49e6-81b1-02bfe409e372@roeck-us.net
Fixes: c9b47e6b2311 ("erofs: cap LZMA stream pool size")
Signed-off-by: Gao Xiang <xiang@kernel.org>
Tested-by: SJ Park <sj@kernel.org>
Tested-by: Geert Uytterhoeven <geert@linux-m68k.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
|
|
exfat_find_free_bitmap() searches the entire allocation bitmap and may
wrap around to its beginning. exfat_trim_fs() does not verify that the
returned cluster is still within the requested FITRIM range.
As a result, a partial FITRIM operation may discard free clusters outside
the user-specified range and report a trimmed length larger than the
requested length.
Validate each returned cluster against the requested range and stop the
search when it wraps around or passes the range end.
Signed-off-by: Yang Wen <anmuxixixi@gmail.com>
Signed-off-by: Namjae Jeon <linkinjeon@kernel.org>
|
|
Sashiko has complained about out of bounds issues if ctx->pos isn't what
is expected in __eventfs_iterate()[1]. This would be an issue if the logic
that calls __eventfs_iterate() didn't already prevent the code from going
out of bounds.
The issue Sashiko brings up is if a user uses lseek64() to put in a
position like 0x100000000 which will overflow the integer used to iterate
the files. This should never be an issue because both tracefs and eventfs
uses the default "maxbytes" for its superblock "s_maxbytes" field which is
defined as:
fs/super.c: s->s_maxbytes = MAX_NON_LFS;
include/linux/fs.h:#define MAX_NON_LFS ((1UL<<31) - 1)
Where MAX_NON_LFS turns into 0x7fffffff.
Testing this with code to try to pass 0x100000000 to lseek64() to a
eventfs directory returns -EINVAL.
But relying on logic for the integrity of a function is not very robust.
Add a WARN_ON_ONCE() in case the ctx->pos is out of the expected range.
[1] https://sashiko.dev/#/patchset/20260810160708.3460a2fd%40gandalf.local.home
Link: https://patch.msgid.link/20260810175928.5f4d9d5c@gandalf.local.home
Signed-off-by: Steven Rostedt <rostedt@goodmis.org>
|
|
In mi_enum_attr(), the start/end VCN validation for non-resident
attributes is:
if (svcn > evcn + 1) goto out;
When evcn is U64_MAX the "evcn + 1" expression wraps to 0 and any svcn
passes the check. For evcn values close to U64_MAX (but not equal to it)
the right-hand side is still a meaningless near-wrap upper bound, so a
malformed on-disk attribute with svcn == 0 and evcn near U64_MAX can pass
mi_enum_attr() unrejected.
VCN (virtual cluster number) is a cluster index, so any valid evcn is
bounded by the volume's total cluster count, which ntfs3 holds in
sbi->used.bitmap.nbits (set up in ntfs_init_from_boot() before any caller
of mi_enum_attr() runs). Reject evcn values that fall outside this range.
However, an empty non-resident attribute (no allocated clusters) is
legitimately encoded with svcn == 0 and evcn == -1 (U64_MAX), e.g. via
attr->nres.evcn = cpu_to_le64((u64)vcn - 1) with vcn == 0. That sentinel
must keep passing, so exclude evcn == U64_MAX from the range check. The
existing "svcn > evcn + 1" test still tolerates the sentinel ("0 > 0" is
false) and continues to require svcn == 0 for it, while the range check
rejects every other out-of-range evcn and thereby also defuses the
"evcn + 1" wraparound.
svcn does not need its own bound: once evcn < nbits, "svcn > evcn + 1"
implies svcn <= nbits.
Fixes: 013ff63b6494 ("fs/ntfs3: Add more attributes checks in mi_enum_attr()")
Signed-off-by: Zhan Xusheng <zhanxusheng@xiaomi.com>
[almaz.alexandrovich@paragon-software.com: fixed evcn check]
Signed-off-by: Konstantin Komarov <almaz.alexandrovich@paragon-software.com>
|
|
In ntfs_init_from_boot(), the boot sector's MFT cluster numbers are
validated against the volume size with:
if (mlcn * sct_per_clst >= sectors ||
mlcn2 * sct_per_clst >= sectors)
goto out;
mlcn and mlcn2 are u64 fields read directly from the boot sector.
sct_per_clst is bounded above by 4096 (true_sectors_per_clst() plus
the is_power_of_2() check below it), but the multiplication is done
in u64 and wraps when mlcn (or mlcn2) is large enough -- e.g. mlcn
near 2^62 with sct_per_clst == 4 wraps to 0, which compares below
any non-zero 'sectors', so the check is bypassed and the malformed
record is accepted.
The accepted mlcn is then used unchanged in
sbi->mft.lbo = mlcn << cluster_bits;
In practice the resulting reads fail at the block layer (sb_bread()
returns NULL via grow_buffers()'s check_mul_overflow() guard), so
today this manifests as mount failing in odd places rather than as
something more dangerous, but the validation step is still wrong
and there is no reason for callers to rely on the block layer to
catch a value that should never have been accepted in the first
place.
Use check_mul_overflow() to compute the two sector positions and
fail the mount if either multiplication wraps; this preserves the
existing semantics (mlcn * sct_per_clst >= sectors) instead of
switching to division (mlcn >= sectors / sct_per_clst), which
would tighten the check at edge cases where 'sectors' is not a
multiple of sct_per_clst. The check_*_overflow() style is the
one ntfs3 already uses for similar on-disk arithmetic in
fs/ntfs3/run.c.
Fixes: 82cae269cfa9 ("fs/ntfs3: Add initialization of super block")
Signed-off-by: Zhan Xusheng <zhanxusheng@xiaomi.com>
Signed-off-by: Konstantin Komarov <almaz.alexandrovich@paragon-software.com>
|
|
If a negative offset is read off disk (for example the offset into the
decompressed fragment block), this will cause squashfs_copy_data() to
perform an out of bounds access.
Fix by checking if offset is negative, and returning 0. This matches
existing behaviour where an offset beyond the block returns 0 bytes
copied.
To trigger this out of bounds access requires a crafted Squashfs
filesystem and CAP_SYS_ADMIN to mount it. Unprivileged users will not be
able to mount such a filesystem, but once mounted, an unprivileged user
can trigger the out of bounds access by reading the crafted file with the
negative offset.
Link: https://lore.kernel.org/20260807162951.672510-1-phillip@squashfs.org.uk
Fixes: f400e12656ab ("Squashfs: cache operations")
Signed-off-by: Phillip Lougher <phillip@squashfs.org.uk>
Reported-by: Yuejie Shi <syjcnss@gmail.com>
Closes: https://lore.kernel.org/all/20260803032735.81785-1-syjcnss@gmail.com/
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
|
|
In ocfs2_dir_foreach_blk_el(), the directory cookie position is
rebuilt with
ctx->pos = (ctx->pos & ~(sb->s_blocksize - 1)) | offset;
`ctx->pos` is loff_t (signed 64-bit), while `sb->s_blocksize` is
unsigned long. On 32-bit kernels unsigned long is 32-bit, so the mask
~(sb->s_blocksize - 1)
is computed as a 32-bit unsigned value (e.g. 0xfffff000 for a 4 KiB
block size). In the AND expression with the 64-bit `ctx->pos`, that
unsigned operand is zero-extended to 64 bits per the usual arithmetic
conversions, yielding 0x00000000fffff000. The high 32 bits of
`ctx->pos` are silently cleared, even though directory size is
allowed to exceed 4 GiB.
When readdir() crosses the 4 GiB boundary on a 32-bit kernel the
position is reset back into the first 4 GiB block, making the
re-validation path re-enumerate already-returned dirents indefinitely.
This is ocfs2_dir_foreach_blk_el(), the extent-list readdir path taken
for all non-inline directories, so a directory large enough to cross
4 GiB reaches it.
This is the same class of bug that commit 3dce5bb82c97 ("exfat: Fix
bitwise operation having different size") fixed in exfat, and the
fix mirrors the equivalent ext4 fix in this series. Cast the operand
to loff_t so the mask is 64-bit before the AND:
ctx->pos = (ctx->pos & ~((loff_t)sb->s_blocksize - 1)) | offset;
64-bit kernels are unaffected.
Link: https://lore.kernel.org/20260806022044.167962-3-zhanxusheng@xiaomi.com
Fixes: ccd979bdbce9 ("[PATCH] OCFS2: The Second Oracle Cluster Filesystem")
Signed-off-by: Zhan Xusheng <zhanxusheng@xiaomi.com>
Reviewed-by: Joseph Qi <joseph.qi@linux.alibaba.com>
Cc: Mark Fasheh <mark@fasheh.com>
Cc: Joel Becker <jlbec@evilplan.org>
Cc: Junxiao Bi <junxiao.bi@oracle.com>
Cc: Changwei Ge <gechangwei@live.cn>
Cc: Jun Piao <piaojun@huawei.com>
Cc: Heming Zhao <heming.zhao@suse.com>
Cc: Andreas Dilger <adilger.kernel@dilger.ca>
Cc: Jan Kara <jack@suse.cz>
Cc: Ojaswin Mujoo <ojaswin@linux.ibm.com>
Cc: "Ritesh Harjani (IBM)" <ritesh.list@gmail.com>
Cc: Ted Ts'o <tytso@mit.edu>
Cc: "zhangyi (F)" <yi.zhang@huawei.com>
Cc: <stable@vger.kernel.org>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
|
|
When reclaiming a suballocator block group, first reduce the on-disk
cluster count by cl_cpg. The current code then subtracts that new count
(fe->i_clusters) from the old cached count
(OCFS2_I(alloc_inode)->ip_clusters).
For an allocator with N block groups, that leaves the cache at
N * cl_cpg - (N * cl_cpg - cl_cpg) = cl_cpg
i.e. ip_clusters -= (fe->i_clusters - cl_cpg) leaves ip_clusters equal to
cl_cpg regardless of N. This happens to be correct when reclaiming from
two block groups, but undercounts the clusters from three block groups
onwards. The incorrect cache value is also used immediately to update
i_blocks.
Assign the updated on-disk count to the cache, matching the allocation and
inode refresh paths.
In a QEMU test using a clean 256 MiB OCFS2 image and a 10,000-file
create/delete workload, the first buggy reclaim left the on-disk
(fe->i_clusters) and cached (ip_clusters) counts at 2048 and 512 clusters
respectively; later reclaims underflowed the cache. With this change, the
cache matched the on-disk count across all four reclaims: 2048, 1536,
1024, and 512 clusters.
Link: https://lore.kernel.org/20260805113920.385959-1-matthias.goergens@gmail.com
Fixes: 4a54331616b3 ("ocfs2: give ocfs2 the ability to reclaim suballocator free bg")
Signed-off-by: Matthias Goergens <matthias.goergens@gmail.com>
Reviewed-by: Joseph Qi <joseph.qi@linux.alibaba.com>
Cc: Mark Fasheh <mark@fasheh.com>
Cc: Joel Becker <jlbec@evilplan.org>
Cc: Junxiao Bi <junxiao.bi@oracle.com>
Cc: Changwei Ge <gechangwei@live.cn>
Cc: Jun Piao <piaojun@huawei.com>
Cc: Heming Zhao <heming.zhao@suse.com>
Cc: <stable@vger.kernel.org>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
|
|
A lockdep warning indicates a circular locking dependency between
`&oi->ip_xattr_sem` and `&journal->j_trans_barrier`:
WARNING: possible circular locking dependency detected
is trying to acquire lock:
(&oi->ip_xattr_sem){++++}-{4:4}, at: ocfs2_init_acl+0x2fd/0x7e0
fs/ocfs2/acl.c:367
but task is already holding lock:
(&journal->j_trans_barrier){.+.+}-{4:4}, at: ocfs2_start_trans+0x3ab/0x700
fs/ocfs2/journal.c:369
The deadlock involves two code paths: Path 1 (setxattr) where
`ocfs2_xattr_set()` acquires `ip_xattr_sem` (write) and then starts a
transaction, which acquires `j_trans_barrier` (read); and Path 2
(mkdir/mknod) where `ocfs2_mknod()` starts a transaction (`j_trans_barrier`
read) and then calls `ocfs2_init_acl()`, which attempts to acquire
`ip_xattr_sem` (read) on the parent directory to retrieve the default ACL.
Because rw_semaphores are subject to writer priority, a pending writer on
`j_trans_barrier` (e.g., the journal commit thread) can cause Path 1 to
block, while Path 2 is blocked waiting for Path 1 to release
`ip_xattr_sem`.
The patch fixes the lock ordering by precomputing the ACL state before
starting the OCFS2 transaction, while preserving POSIX ACL storage
semantics and the existing inode/security initialization order. By reading
the parent directory's default ACL and preparing the new inode's ACLs
outside the transaction, `ip_xattr_sem` is always acquired before
`j_trans_barrier`.
`struct ocfs2_acl_state` encapsulates the prepared ACL state, while
`ocfs2_acl_init_prepare()` and `ocfs2_acl_init_release()` avoid code
duplication between `ocfs2_mknod()` and `ocfs2_init_security_and_acl()`.
`ocfs2_calc_xattr_init()` and `ocfs2_init_acl()` use this precomputed
state, removing internal `ip_xattr_sem` acquisition and redundant disk
reads.
Additionally, remove the `ip_xattr_sem` acquisition from
`ocfs2_xattr_set_handle()`. This function is only used while initializing a
new inode that has not yet been inserted into the inode hash or attached to
a dentry, meaning there is no risk of concurrent access and the lock is
unnecessary.
Link: https://lore.kernel.org/4094de06-9b69-4174-b2ee-08126dffc693@mail.kernel.org
Fixes: 16c8d569f570 ("ocfs2/acl: use 'ip_xattr_sem' to protect getting extended attribute")
Signed-off-by: Krystian Kaniewski <krystianmkaniewski@gmail.com>
Assisted-by: Gemini:gemini-3.5-flash Gemini:gemini-3.1-pro-preview syzbot
Reported-by: syzbot+4007ab5229e732466d9f@syzkaller.appspotmail.com
Closes: https://syzkaller.appspot.com/bug?extid=4007ab5229e732466d9f
Link: https://syzkaller.appspot.com/ai_job?id=cc75363d-c672-499e-8fc5-44bcdc1cee39
Reviewed-by: Joseph Qi <joseph.qi@linux.alibaba.com>
Cc: Mark Fasheh <mark@fasheh.com>
Cc: Joel Becker <jlbec@evilplan.org>
Cc: Junxiao Bi <junxiao.bi@oracle.com>
Cc: Changwei Ge <gechangwei@live.cn>
Cc: Jun Piao <piaojun@huawei.com>
Cc: Heming Zhao <heming.zhao@suse.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
|
|
[BUG]
A corrupted append-DIO dinode (high byte at offset 0xa1
corrupted from 0 to 1) can carry an i_dio_orphaned_slot
outside the mounted filesystem slot range and trigger a
use-after-free error:
BUG: KASAN: slab-use-after-free in ocfs2_get_system_file_inode+0x780/0x820 fs/ocfs2/sysfile.c:102
Read of size 8 at addr ffff88800b767c00 by task kworker/u8:3/85
Call Trace:
...
ocfs2_get_system_file_inode+0x780/0x820 fs/ocfs2/sysfile.c:102
ocfs2_wipe_inode+0x292/0xf70 fs/ocfs2/inode.c:840
ocfs2_delete_inode fs/ocfs2/inode.c:1155 [inline]
ocfs2_evict_inode+0x6c9/0x1170 fs/ocfs2/inode.c:1295
evict+0x38e/0x8f0 fs/inode.c:810
iput_final fs/inode.c:1914 [inline]
iput fs/inode.c:1966 [inline]
iput+0x55b/0x8b0 fs/inode.c:1926
ocfs2_recover_orphans+0x610/0xe40 fs/ocfs2/journal.c:2374
ocfs2_complete_recovery+0x5af/0xd00 fs/ocfs2/journal.c:1373
...
[CAUSE]
ocfs2_del_inode_from_orphan() uses i_dio_orphaned_slot to index the
slot-local system inode cache. The dinode validator does not check
this active slot, so an out-of-range value produces an invalid cache
entry pointer that is dereferenced as an inode pointer.
[FIX]
Reject an active i_dio_orphaned_slot outside the slot range during
dinode validation, before DIO orphan recovery can consume it.
Link: https://lore.kernel.org/20260803030007.3993199-3-gality369@gmail.com
Fixes: 06ee5c75b575 ("ocfs2: add functions to add and remove inode in orphan dir")
Signed-off-by: ZhengYuan Huang <gality369@gmail.com>
Reviewed-by: Joseph Qi <joseph.qi@linux.alibaba.com>
Cc: Mark Fasheh <mark@fasheh.com>
Cc: Joel Becker <jlbec@evilplan.org>
Cc: Junxiao Bi <junxiao.bi@oracle.com>
Cc: Changwei Ge <gechangwei@live.cn>
Cc: Jun Piao <piaojun@huawei.com>
Cc: Heming Zhao <heming.zhao@suse.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
|
|
Patch series "ocfs2: validate active orphan slots during inode read".
OCFS2 trusts active ordinary and append-DIO orphan slots read from dinodes.
A corrupted slot can therefore index osb_orphan_wipes or the slot-local
system-inode cache outside their allocations before the corruption is
reported.
Patch 1 validates the ordinary orphan slot used by inode wipe processing.
Patch 2 validates the append-DIO orphan slot used by DIO completion and
orphan recovery. Both checks reject corrupt metadata at the existing inode
validation boundary.
This patch (of 2):
[BUG]
A corrupted dinode with OCFS2_ORPHANED_FL can carry an
i_orphaned_slot outside the mounted filesystem slot range.
ocfs2_wipe_inode() uses it to index osb_orphan_wipes before looking
up the orphan directory, causing an out-of-bounds memory access.
BUG: KASAN: slab-use-after-free in ocfs2_get_system_file_inode+0x780/0x820 fs/ocfs2/sysfile.c:102
Read of size 8 at addr ffff88800b767c00 by task kworker/u8:3/85
Call Trace:
...
ocfs2_get_system_file_inode+0x780/0x820 fs/ocfs2/sysfile.c:102
ocfs2_wipe_inode+0x292/0xf70 fs/ocfs2/inode.c:840
ocfs2_delete_inode fs/ocfs2/inode.c:1155 [inline]
ocfs2_evict_inode+0x6c9/0x1170 fs/ocfs2/inode.c:1295
evict+0x38e/0x8f0 fs/inode.c:810
iput_final fs/inode.c:1914 [inline]
iput fs/inode.c:1966 [inline]
iput+0x55b/0x8b0 fs/inode.c:1926
ocfs2_recover_orphans+0x610/0xe40 fs/ocfs2/journal.c:2374
ocfs2_complete_recovery+0x5af/0xd00 fs/ocfs2/journal.c:1373
...
[CAUSE]
ocfs2_validate_inode_block() validates i_suballoc_slot but leaves
the active ordinary orphan slot unchecked. Downstream consumers
assume that the value is smaller than osb->max_slots.
[FIX]
Reject an active i_orphaned_slot outside the slot range during
dinode validation, before the inode reaches orphan wipe processing.
Link: https://lore.kernel.org/20260803030007.3993199-1-gality369@gmail.com
Link: https://lore.kernel.org/20260803030007.3993199-2-gality369@gmail.com
Fixes: b4df6ed8db0c ("[PATCH] ocfs2: fix orphan recovery deadlock")
Signed-off-by: ZhengYuan Huang <gality369@gmail.com>
Reviewed-by: Joseph Qi <joseph.qi@linux.alibaba.com>
Cc: Mark Fasheh <mark@fasheh.com>
Cc: Joel Becker <jlbec@evilplan.org>
Cc: Junxiao Bi <junxiao.bi@oracle.com>
Cc: Changwei Ge <gechangwei@live.cn>
Cc: Jun Piao <piaojun@huawei.com>
Cc: Heming Zhao <heming.zhao@suse.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
|
|
Clear the local crypto-related structures via __cleanup() functions
when we're done with them to avoid that sensitive data could leak on
the stack.
Note: cmac_ctx in ksmbd_sign_smb3_pdu() gets cleared in aes_cmac_final()
already, so this does not need a __cleanup() marker.
Signed-off-by: Thomas Huth <thuth@redhat.com>
Acked-by: Namjae Jeon <linkinjeon@kernel.org>
Link: https://patch.msgid.link/20260807125845.1477067-3-thuth@redhat.com
Signed-off-by: Eric Biggers <ebiggers@kernel.org>
|
|
- Call f2fs_add_ino_entry() and f2fs_remove_ino_entry() for ORPHAN_INO
- introduce __f2fs_add_ino_entry() to wrap __add_ino_entry(), so that
both f2fs_add_ino_entry() and f2fs_set_dirty_device() will call
__f2fs_add_ino_entry().
So, after this change:
add delete lookup
ORPHAN_INO f2fs_add_ino_entry f2fs_remove_ino_entry N/A
FLUSH_INO f2fs_set_dirty_device f2fs_remove_ino_entry f2fs_is_dirty_device
APPEND_INO f2fs_add_ino_entry f2fs_remove_ino_entry f2fs_exist_written_data
UPDATA_INO f2fs_add_ino_entry f2fs_remove_ino_entry f2fs_exist_written_data
TRANS_DIR_INO f2fs_add_ino_entry N/A f2fs_exist_written_data
XATTR_DIR_INO f2fs_add_ino_entry N/A f2fs_exist_written_data
Signed-off-by: Chao Yu <chao@kernel.org>
Signed-off-by: Jaegeuk Kim <jaegeuk@kernel.org>
|
|
generic/794 4s ... - output mismatch (see /share/git/fstests/results//generic/794.out.bad)
--- tests/generic/794.out 2026-06-12 08:46:32.766426241 +0800
+++ /share/git/fstests/results//generic/794.out.bad 2026-07-05 18:32:55.000000000 +0800
@@ -1,4 +1,16 @@
QA output created by 794
append_write
+FAIL: non-zero data in gap [4080,4096) after shutdown+remount
+000000 5a 5a 5a 5a 5a 5a 5a 5a 5a 5a 5a 5a 5a 5a 5a 5a >ZZZZZZZZZZZZZZZZ<
+*
+001000
truncate_up
...
(Run 'diff -u /share/git/fstests/tests/generic/794.out /share/git/fstests/results//generic/794.out.bad' to see the entire diff)
Ran: generic/794
Failures: generic/794
Failed 1 of 1 tests
Steps of generic/794:
1. write 4096 bytes to file w/ 0x5a
2. use fiemap to get PBA of first block in file
3. truncate file to 4080
4. umount; write 4096 bytes to file w/ 0x5a directly via PBA; mount
5. extend filesize via
a) append 4096 from offset 4096, or
b) truncate 8192, or
c) fallocate 4096 from offset 4096
6. verify the gap is zeroed in memory [4080,4096)
7. sync range 4096 from offset 4096; shutdown -f (flush meta before shutdown)
8. umount; mount; verify [4080,4096) is zeroed or not.
When extending file size (e.g. via truncate, fallocate, or write) across an
unaligned EOF boundary, we need to ensure that post-EOF data in the partial
page is zeroed out in pagecache and marked dirty, then writeback the cache to
persist zeroed data before committing inode w/ updated i_size.
This help to prevent stale disk data beyond the previous EOF from being exposed
after remounting or crash recovery.
Since f2fs is a LFS filesystem, we only support direct write via PBA in pinfile,
and pinfile has section-aligned filesize, so in Android, there should no problem,
but for other usage in different environment, let's fix this w/ fsync_mode=strict
mount option.
Cc: stable@kernel.org
Signed-off-by: Chao Yu <chao@kernel.org>
Signed-off-by: Jaegeuk Kim <jaegeuk@kernel.org>
|
|
Otherwise, it will drop one more page after new_size which is not
necessary.
Cc: stable@kernel.org
Fixes: ba8dac350faf ("f2fs: fix to zero post-eof page")
Signed-off-by: Chao Yu <chao@kernel.org>
Signed-off-by: Jaegeuk Kim <jaegeuk@kernel.org>
|
|
ceph_ioctl_set_layout() and ceph_ioctl_set_layout_policy() call
inode_owner_or_capable() with &nop_mnt_idmap instead of the idmap of the
mount the ioctl was issued on.
CephFS supports idmapped mounts (FS_ALLOW_IDMAP), so on such a mount this
compares the caller's fsuid against the unmapped on-disk owner rather than
the mapped owner: the actual owner can be wrongly denied with -EACCES and
an unrelated caller wrongly allowed. Both functions already have the
struct file, so use file_mnt_idmap(file) instead.
Cc: stable@vger.kernel.org
Fixes: cee38bbf5556 ("ceph: add owner/capability checks for CEPH_IOC_SET_LAYOUT*")
Signed-off-by: Zhan Xusheng <zhanxusheng@xiaomi.com>
Reviewed-by: Xiubo Li <xiubo.li@clyso.com>
Reviewed-by: Alex Markuze <amarkuze@redhat.com>
Signed-off-by: Ilya Dryomov <idryomov@gmail.com>
|
|
CEPH_MDS_IS_READY() is parsed so that the ternary expression can
return true for an MDS entry with state 0 when it is not laggy. This
allows the random selector to choose a down/DNE rank.
Group the ternary expression under the state check so zero-state ranks
are not treated as ready.
Cc: stable@vger.kernel.org
Fixes: b38c9eb4757d ("ceph: add possible_max_rank and make the code more readable")
Link: https://tracker.ceph.com/issues/78648
Signed-off-by: Yiming Zhu <zhuyiming@kuaishou.com>
Reviewed-by: Viacheslav Dubeyko <slava@dubeyko.com>
Signed-off-by: Ilya Dryomov <idryomov@gmail.com>
|
|
Function gfs2_glock_hold() is expected only to be called when the glock is
held, so use lockref_get_not_zero() instead of lockref_get_not_dead().
In addition, when an asynchronous callback arrives in gfs2_glock_cb(), the
glock can already be dead (from __gfs2_glock_put()), or it can be on the glock
lru list with refcount 0, so we cannot use gfs2_glock_hold() there.
With gfs2_glock_cb() now handling dead glocks, we can remove the racy check in
gdlm_bast().
Signed-off-by: Andreas Gruenbacher <agruenba@redhat.com>
|
|
It was an effort to enhance fscache as a kernel cache for lazy
pulling (at least according to previous Incremental FS discussion [1])
and EROFS over fscache was the in-tree user of this mode.
fscache has since evolved to be netfslib-oriented, serving network
filesystem inodes via the netfs library, but EROFS never acts as a
network filesystem and we need to cache golden filesystem images rather
than individual EROFS inodes.
Since EROFS over fscache is now removed, clean up netfs/fscache/
cachefiles upstream too.
[1] https://lore.kernel.org/r/CAOQ4uxi4dzxArY24YO=+kBCK2gGoq3Ptb8WkzCqSogPgU_R3dQ@mail.gmail.com
[dh] Fixed up comments on:
https://sashiko.dev/#/patchset/20260716103030.3065561-1-dhowells%40redhat.com
https://sashiko.dev/#/patchset/20260722130218.78958-1-dhowells%40redhat.com
Signed-off-by: Gao Xiang <xiang@kernel.org>
Signed-off-by: David Howells <dhowells@redhat.com>
Link: https://patch.msgid.link/1046393.1786544127@warthog.procyon.org.uk
cc: Paulo Alcantara <pc@manguebit.org>
cc: netfs@lists.linux.dev
cc: linux-erofs@lists.ozlabs.org
cc: bpf@vger.kernel.org
cc: linux-fsdevel@vger.kernel.org
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
|
|
ovl_create_real() releases the new dentry twice when the casefold
consistency check fails. The S_IFDIR branch calls end_creating() and
sets err, then falls through to the common out: label which calls
end_creating() on the same dentry again:
case S_IFDIR:
newdentry = ovl_do_mkdir(ofs, dir, newdentry, attr->mode);
err = PTR_ERR_OR_ZERO(newdentry);
if (!err && ofs->casefold != ovl_dentry_casefolded(newdentry)) {
pr_warn_ratelimited(...);
end_creating(newdentry); /* first */
err = -EINVAL;
}
break;
...
if (err)
goto out;
...
out:
if (err) {
end_creating(newdentry); /* second, same dentry */
return ERR_PTR(err);
}
end_creating() is end_dirop(), which does inode_unlock() on the parent
and dput() on the dentry, so the parent directory's i_rwsem is unlocked
twice and the dentry is put twice. The second unlock releases a lock
that is not held, which is what wedges every later creation under that
parent, and the second dput() drops a reference that was never taken.
The branch was added by commit dfc7da402ccc ("ovl: Check for casefold
consistency when creating new dentries") as a bare dput(), which already
released the reference twice; commit fe497f0759e0 ("VFS: change
vfs_mkdir() to unlock on failure.") converted both sites to
end_creating(), adding the double unlock.
This is reachable by an unprivileged user. The casefold consistency of
the layers is validated at mount time in ovl_parse_layer(), and again on
every lookup in ovl_lookup_single(), but ofs->workdir is the internal
"work" subdirectory created inside the user-supplied workdir, and that
subdirectory is not re-checked. Marking it casefolded after the mount
therefore makes every ovl_create_temp() inherit the wrong state - and
that path reaches ovl_create_real() through ovl_start_creating_temp(),
which uses start_creating() with a generated name and so never runs the
lookup-time check.
unshare -Urm
mount -t tmpfs -o casefold=utf8-12.1.0 tmpfs mnt
mkdir -p mnt/lower/d mnt/upper mnt/work mnt/merged
mount -t overlay ovl -o lowerdir=mnt/lower,\
upperdir=mnt/upper,workdir=mnt/work mnt/merged
chattr +F mnt/work/work
mkdir mnt/merged/d/sub # directory copy-up
overlayfs: wrong inherited casefold (work/#5)
and the next copy-up blocks forever on the parent's i_rwsem:
mkdir D start_creating+0x65/0xb0
ovl_start_creating_temp+0xb0/0xe0 [overlay]
ovl_create_temp+0xa3/0x1d0 [overlay]
ovl_copy_up_one+0x1f1c/0x21c0 [overlay]
ovl_copy_up_flags+0xf5/0x140 [overlay]
ovl_create_object+0xb7/0x220 [overlay]
ovl_mkdir+0x23/0x40 [overlay]
Drop the end_creating() from the branch and let out: own the cleanup,
which is what every other error path in this function already does.
Fixes: dfc7da402ccc ("ovl: Check for casefold consistency when creating new dentries")
Cc: stable@vger.kernel.org
Signed-off-by: Vivek Parikh <vivek.parikh@breachx.ai>
Reviewed-by: Amir Goldstein <amir73il@gmail.com>
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
|
|
pipe_poll() unconditionally sets ->poll_usage on the first call, forcing
anon_pipe_write() to wake up readers on every write even if the pipe was
not empty.
The reason is that some legacy epoll(EPOLLET) users depend on historical
per-write wakeups, see commit 3a34b13a88ca ("pipe: make pipe writes always
wake up readers").
Test-case:
#include <unistd.h>
#include <sys/epoll.h>
#include <assert.h>
int main(void)
{
int pfd[2], efd;
struct epoll_event evt = { .events = EPOLLIN | EPOLLET };
pipe(pfd);
efd = epoll_create1(0);
epoll_ctl(efd, EPOLL_CTL_ADD, pfd[0], &evt);
for (int i = 0; i < 2; ++i) {
write(pfd[1], "", 1);
assert(epoll_wait(efd, &evt, 1, 0) == 1);
}
return 0;
}
it fails if WRITE_ONCE(poll_usage, true) is removed from pipe_poll().
However, without EPOLLET in .events, it does not need the extra wakeup
and succeeds even if write() is called only once before the main loop.
Currently io_uring without (unsupported) IORING_POLL_ADD_LEVEL always
sets EPOLLET, and in IORING_POLL_ADD_MULTI mode it depends on per-write
wakeups the same way:
#include <unistd.h>
#include <sys/mman.h>
#include <sys/epoll.h>
#include <sys/syscall.h>
#include <linux/io_uring.h>
#include <assert.h>
int main(void)
{
struct io_uring_params p = {};
int fd, pfd[2];
pipe(pfd);
fd = syscall(SYS_io_uring_setup, 2, &p);
assert(fd >= 0);
void *ring = mmap(0, p.cq_off.cqes + p.cq_entries * sizeof(struct io_uring_cqe),
PROT_READ | PROT_WRITE, MAP_SHARED, fd, IORING_OFF_SQ_RING);
assert(ring != MAP_FAILED);
*(unsigned *)(ring + p.sq_off.tail) = 1;
struct io_uring_sqe *sqes = mmap(0, p.sq_entries * sizeof(*sqes),
PROT_READ | PROT_WRITE, MAP_SHARED, fd, IORING_OFF_SQES);
assert(sqes != MAP_FAILED);
sqes[0].opcode = IORING_OP_POLL_ADD;
sqes[0].fd = pfd[0];
sqes[0].len = IORING_POLL_ADD_MULTI;
sqes[0].poll32_events = EPOLLIN;
syscall(SYS_io_uring_enter, fd, 1, 0, 0, 0, 0);
unsigned *cq_head = ring + p.cq_off.head;
unsigned *cq_tail = ring + p.cq_off.tail;
for (int i = 0; i < 2; ++i) {
write(pfd[1], "", 1);
syscall(SYS_io_uring_enter, fd, 0, 0, IORING_ENTER_GETEVENTS, 0, 0);
assert(*cq_tail == ++*cq_head);
}
return 0;
}
the 2nd assert() in the main loop fails without ->poll_usage == true.
Rename ->poll_usage to ->pseudo_edgetrigger to make the purpose clearer,
update the comments, and change pipe_poll() to set ->pseudo_edgetrigger
only if wait->_key & EPOLLET is true. This check should catch both users,
and this way poll/select and epoll without EPOLLET users will not pay for
the extra wakeup.
Signed-off-by: Oleg Nesterov <oleg@redhat.com>
Link: https://patch.msgid.link/anCNoW-x0bcB2ggg@redhat.com
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
|
|
The PIDFD_GET_*_NAMESPACE ioctls in pidfd_ioctl() perform a filesystem
credentials ptrace access check before handing out a namespace file
descriptor. The accompanying comment states that the code "mirrors nsfs
behavior", but, unlike the corresponding procfs paths, it does so without
holding the target task's exec_update_lock.
proc_ns_get_link() and proc_ns_readlink() both take exec_update_lock for
reading around the ptrace check and the namespace lookup, so that the
credentials used for the access decision match those of the task when its
namespace is read. Without it, a caller can pass the check against the
target's old credentials and then read the namespace after the target has
execve()'d a setuid binary and committed new credentials -- accessing
namespace information it should have been denied.
Hold exec_update_lock for reading around the ptrace check and the
namespace lookup so that pidfd truly mirrors nsfs behavior, as the comment
already claims. open_namespace() itself runs outside the lock: once a
namespace reference is obtained it carries its own refcount and is opened
with the caller's own credentials, so a concurrent execve() on the target
can no longer affect the outcome.
Fixes: 5b08bd408534 ("pidfs: allow retrieval of namespace file descriptors")
Cc: stable@vger.kernel.org
Signed-off-by: Chen Linxuan <me@black-desk.cn>
Link: https://patch.msgid.link/20260731-pidfd-exec-update-lock-v1-1-b388f2f3a8b0@black-desk.cn
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
|
|
backing_file_open() derives the path to be stored in the new backing
file from user_file->f_path. This is incorrect when user_file itself
is a backing file, which is the case for nested stacking filesystems,
e.g. overlayfs mounts where the lowerdir of one overlayfs is the merged
directory of another. Since commit def3ae83da02 ("fs: store real path
instead of fake path in backing file f_path") the f_path of a backing
file holds the real path of the intermediate layer, not the path that
the user opened.
Commit 924577e4f6ca ("ovl: Fix nested backing file paths") fixed this
for such configurations by passing file_user_path() from
ovl_open_realfile(). However, commit 6af36aeb147a ("lsm: add
backing_file LSM hooks") changed the first argument of
backing_file_open() from the user path back to the user file and
derived the path from user_file->f_path again, silently re-introducing
the problem.
As a result, files mapped through a nested overlayfs show the wrong
path in /proc/<pid>/maps and in perf/ftrace mmap records. For example,
with two nested overlayfs mounts:
mkdir -p /ovl/{lower,upper,work,merged} /ovl/nested
echo hello > /ovl/lower/foo
mount -t overlay overlay \
-o lowerdir=/ovl/lower,upperdir=/ovl/upper,workdir=/ovl/work \
/ovl/merged
# at least two lowerdirs are needed when upperdir is nonexistent
mount -t overlay overlay \
-o lowerdir=/ovl/merged:/ovl/lower /ovl/nested
mapping /ovl/nested/foo shows a disconnected path instead of the user
path:
# readlink /proc/self/fd/3
/ovl/nested/foo
# grep foo /proc/self/maps
7f6e2c100000-7f6e2c101000 r--s 00000000 00:24 15813027 /foo
The bogus path is derived from the f_path of the intermediate backing
file, whose mount is a private clone that d_path() cannot resolve.
Fix this by using file_user_path(), which returns the outermost
user-visible path for backing files and falls back to
&user_file->f_path for regular files. This restores the behavior of
commit 924577e4f6ca ("ovl: Fix nested backing file paths") for
overlayfs and also fixes the same problem for the other
backing_file_open() callers, fuse passthrough and erofs ishare, when
their user file is itself a backing file.
backing_tmpfile_open() has the same pattern but is not affected: it is
only called by ovl_create_tmpfile() for the upper layer, and another
overlayfs is rejected as upperdir by the DCACHE_OP_REAL check in
ovl_mount_dir_check(), so its user_file can never be a backing file.
Fixes: 6af36aeb147a ("lsm: add backing_file LSM hooks")
Cc: stable@vger.kernel.org
Signed-off-by: Baokun Li <libaokun@linux.alibaba.com>
Link: https://patch.msgid.link/20260804034204.3487077-1-libaokun@linux.alibaba.com
Tested-by: Paul Moore <paul@paul-moore.com>
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
|
|
inode_insert5() no longer has an isnew argument, but its
kernel-doc still documents one. This triggers a W=1 kernel-doc
warning.
Remove the stale parameter description.
Signed-off-by: Yichong Chen <chenyichong@uniontech.com>
Link: https://patch.msgid.link/20260805024149.935769-1-chenyichong@uniontech.com
Reviewed-by: Jan Kara <jack@suse.cz>
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
|
|
The case labels in the sysfs(2) syscall implementation are indented
one level deeper than the switch statement itself, which does not
match the kernel coding style (switch and case should be at the same
indentation level). Fix the indentation; no functional change.
Signed-off-by: Manush Prajwal <manushprajwal555@gmail.com>
Link: https://patch.msgid.link/20260808182816.2399-1-manushprajwal555@gmail.com
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
|
|
Which could leak blkg references.
Fixes: c03cea42149d ("iomap: add initial support for writes without buffer heads")
Signed-off-by: Christoph Hellwig <hch@lst.de>
Link: https://patch.msgid.link/20260804124404.737145-3-hch@lst.de
Reviewed-by: "Darrick J. Wong" <djwong@kernel.org>
Reviewed-by: Anuj Gupta <anuj20.g@samsung.com>
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
|
|
fs_bio_integrity_alloc might not allocate a bio integrity payload if PI
verification is disabled on the block device. Check for that case before
calling fs_bio_integrity_free in iomap_bio_read_folio_range_sync to
avoid a NULL pointer dereferences.
Make the branch cover the PI verification as well - while
fs_bio_integrity_verify works without an integrity payload, it requires
one to actually do useful work.
Fixes: 0b10a370529c ("iomap: support T10 protection information")
Cc: stable@vger.kernel.org # v7.1
Signed-off-by: Christoph Hellwig <hch@lst.de>
Reviewed-by: Anuj Gupta <anuj20.g@samsung.com>
Reviewed-by: Kanchan Joshi <joshi.k@samsung.com>
Reviewed-by: "Darrick J. Wong" <djwong@kernel.org>
Link: https://patch.msgid.link/20260804124404.737145-2-hch@lst.de
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
|
|
WARN_ON_ONCE() takes a condition, not a message. The string literals
are always true, so the warnings still trigger but the messages are
never printed.
Use WARN_ONCE(1, ...) instead to print the messages and keep the
once-only behavior.
Found with a Coccinelle script. Clang's -Wstring-conversion also flags
such calls but is not enabled in kernel builds.
Fixes: f0cd988016f6 ("fs: massage locking helpers")
Signed-off-by: Karl Mehltretter <kmehltretter@gmail.com>
Link: https://patch.msgid.link/20260808123802.73687-1-kmehltretter@gmail.com
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
|
|
__hfsplus_ext_write_extent() writes the cached extent record back into a
B-tree node using fd->entrylength as the length, and fd->entrylength is
derived in __hfs_brec_find() from two on-disk values:
fd->entrylength = len - keylen;
A crafted image can keep both len and keylen valid but make fd->entrylength
negative (keylen > len). __hfsplus_ext_write_extent() doesn't check
fd->entrylength before consuming it, and hfs_bnode_write() takes the length
as u32, so the negative value turns into a huge one. The copy then reads
data past the end of hip->cached_extents, which is only
sizeof(hfsplus_extent_rec) bytes long, and leaks kernel memory into the
image.
Reject an fd->entrylength that does not match sizeof(hfsplus_extent_rec) in
__hfsplus_ext_write_extent(), mirroring the check already performed in
__hfsplus_ext_read_extent().
Link: https://lore.kernel.org/lkml/cbd7003314c530d4f910eacf019ff80adad6687e.camel@dubeyko.com/
Signed-off-by: Jiaming Zhang <r772577952@gmail.com>
Reviewed-by: Viacheslav Dubeyko <slava@dubeyko.com>
Link: https://lore.kernel.org/r/20260810092422.1691377-1-r772577952@gmail.com
Signed-off-by: Viacheslav Dubeyko <slava@dubeyko.com>
|
|
A crafted HFS+ image can contain a corrupted B-tree node. The node
descriptor may contain a record count that does not fit in the node, and
record offsets may be unordered, unaligned, outside the node, or point into
the offset table itself.
Several B-tree helpers consume these on-disk fields before validating them:
hfs_bnode_dump() can walk past the offset table when num_recs is corrupted,
hfs_brec_lenoff() can produce an underflowed length or a record range that
overlaps the offset table. This can make the unlink/writeback path
repeatedly call hfs_bnode_read_u16() with invalid offsets while holding the
HFS+ B-tree lock, producing a flood of "requested invalid offset" messages.
Other writeback workers then block on tree->tree_lock and the system
reports tasks hung in hfsplus_write_inode().
Validate num_recs against the node size before walking the record offset
table. Reject record ranges that are unordered, unaligned, outside the
node, or overlapping the offset table. Reject invalid record indexes before
reading their offset entries, and avoid decrementing an already-zero
leaf_count.
Closes: https://lore.kernel.org/lkml/CANypQFb_2TqKGrztAXj5m0_v+QChxXDnQVeifzV8J25Vuju10Q@mail.gmail.com/
Assisted-by: Codex:gpt-5.5-xhigh
Signed-off-by: Jiaming Zhang <r772577952@gmail.com>
Reviewed-by: Viacheslav Dubeyko <slava@dubeyko.com>
Tested-by: Viacheslav Dubeyko <slava@dubeyko.com>
Signed-off-by: Viacheslav Dubeyko <slava@dubeyko.com>
Link: https://lore.kernel.org/r/20260806073358.1184938-1-r772577952@gmail.com
Signed-off-by: Viacheslav Dubeyko <slava@dubeyko.com>
|