summaryrefslogtreecommitdiff
AgeCommit message (Collapse)AuthorFilesLines
2026-08-06resolve_btfids: Deduplicate BTF after btf2btf transformationsIhor Solodrai1-0/+6
btf2btf() adds new types to the BTF: the KF_IMPLICIT_ARGS transform synthesizes an _impl FUNC together with its FUNC_PROTO and copies of the kfunc's decl tags. Nothing deduplicates them afterwards. pahole runs btf__dedup() on its own output, but that happens before resolve_btfids sees the BTF, so any type the tool itself creates is emitted as-is, even when a structurally identical type is already present. Call btf__dedup() at the start of finalize_btf(), so that base distillation and the by-name sort both operate on the canonical set of types. On an x86_64 build with the BPF selftests config this removes 17 duplicate FUNC_PROTOs from vmlinux BTF. The dedup call increases runtime of resolve_btfids on vmlinux by 30-40%. The performance hit is an acceptable cost to keep kernel BTF deduped [1]. [1] https://lore.kernel.org/bpf/986e6f4e-4b51-4440-a37c-9624906d7370@linux.dev/ Signed-off-by: Ihor Solodrai <ihor.solodrai@linux.dev> Link: https://patch.msgid.link/20260807032029.78092-2-ihor.solodrai@linux.dev Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
2026-08-07bus: fsl-mc: drop unused assignment of acpi_device_id::driver_dataPawel Zalewski (The Capable Hub)1-1/+1
This module sets the acpi_device_id::driver_data to 0 but the field is not actually used within the module, we can just drop it from the table. While we are at it - use a named initializer for the acpi_device_id::id field as well to make the code more readable. Signed-off-by: Pawel Zalewski (The Capable Hub) <pzalewski@thegoodpenguin.co.uk> Reviewed-by: Ioana Ciornei <ioana.ciornei@nxp.com> Link: https://lore.kernel.org/r/20260728-acpi-bus-v1-1-12ff25fdea9b@thegoodpenguin.co.uk Signed-off-by: Christophe Leroy (CS GROUP) <chleroy@kernel.org>
2026-08-07soc: fsl: qe: check platform_driver_register() in qe_ic_of_init()Linkai Gong1-2/+1
qe_ic_of_init() ignored the return value of platform_driver_register() and always returned success. Propagate the error to the initcall. Fixes: be7ecbd240b2 ("soc: fsl: qe: convert QE interrupt controller to platform_device") Signed-off-by: Linkai Gong <gonglinkai@kylinos.cn> Reviewed-by: Maxim Kochetkov <fido_max@inbox.ru> Link: https://lore.kernel.org/r/20260731094608.1883391-1-gonglinkai@kylinos.cn Signed-off-by: Christophe Leroy (CS GROUP) <chleroy@kernel.org>
2026-08-07phy: lynx-10g: use RCW override procedure for dynamic protocol changeVladimir Oltean2-9/+15
Up until this patch, the only protocol change supported was between 1000Base-X/SGMII and 2500Base-X. The others require an RCW override procedure which was lacking. Since now the guts driver provides the means of applying this procedure, make use of it and remove any comment which mentioned the limitation. Signed-off-by: Vladimir Oltean <vladimir.oltean@nxp.com> Acked-by: Vinod Koul <vkoul@kernel.org> Reviewed-by: Ioana Ciornei <ioana.ciornei@nxp.com> Link: https://lore.kernel.org/r/20260721231603.67865-10-vladimir.oltean@nxp.com Signed-off-by: Christophe Leroy (CS GROUP) <chleroy@kernel.org>
2026-08-07soc: fsl: guts: implement the RCW override procedureVladimir Oltean2-8/+409
Add support for the RCW override procedure which enables runtime reconfiguration of the protocol running on a SerDes lane. The procedure is done through the DCFG DCSR space which now can be defined as the second memory region of the guts DT node. Support is added on the following SoCs: LS1046A, LS1088A, LS2088A. The procedure is exported to the "client" driver - the Lynx10G SerDes PHY driver - through the following functions: - fsl_guts_lane_validate() used to validate that changing the protocol on a specific lane is supported. - fsl_guts_lane_set_mode() which can be used to request the RCW procedure be executed for a specific lane. Since the RCW override procedure is different depending on the SoC, the private fsl_soc_data structure is updated with two new per SoC callbacks (.serdes_get_rcw_override() and .serdes_init_rcwcr()) which get used from the generic fsl_guts_lane_set_mode() function. These two callbacks hide all the SoC specific register offsets, masks and values so that the _set_mode() procedure is straightforward. Signed-off-by: Ioana Ciornei <ioana.ciornei@nxp.com> Signed-off-by: Vladimir Oltean <vladimir.oltean@nxp.com> Reviewed-by: Ioana Ciornei <ioana.ciornei@nxp.com> Link: https://lore.kernel.org/r/20260721231603.67865-9-vladimir.oltean@nxp.com Signed-off-by: Christophe Leroy (CS GROUP) <chleroy@kernel.org>
2026-08-07dt-bindings: fsl: layerscape-dcfg: define DCFG_DCSR regionVladimir Oltean1-1/+14
In Layerscape (Arm) and QorIQ (PowerPC) devices, hardware peripherals are accessed by the CPU through a portion of the SoC address space called CCSR ("Configuration, Control, and Status Registers"). All hardware IP blocks have their registers mapped here, and the Device Configuration block makes no exception. However, there exists a secondary range of the address space named DCSR ("Debug Control and Status Registers") which, like CCSR, also holds registers of hardware IP blocks, except the DCSR contents is hidden in all public reference manuals. The intention of the CCSR/DCSR split, to the best of my knowledge, was to place the functionality that is too low level for normal use, and which is necessary only for debug, in a completely separate address space which can be hidden. A use case has appeared where networking SerDes lanes need to be reconfigured at runtime for a different protocol (example: 10GBase-R to SGMII), and the architecture of the SoCs does not normally permit that. The Reset Configuration Word (RCW) is a data structure read by the SoC preboot loader (PBL) which contains stuff like pinmuxing and SerDes protocol mapping for each lane. The RCW that the PBL has loaded is visible in the DCFG block's normal status registers (from CCSR), as read only. Turns out, the RCW is also mapped in the DCFG's shadow register map (in DCSR), in a write-only form. Writing to the RCW registers from the DCFG's DCSR space to change what the PBL has loaded is called "RCW override". It has been validated that the RCW override procedure is necessary to reconfigure the networking data path when a SerDes lane performs a major protocol change. It changes some internal muxes which connect the PCS to either the 10G MAC or to the 1G MAC. Defining the DCSR area of the DCFG as a secondary 'reg' array element allows operating systems to perform RCW overrides. Since it is introduced late in the binding's lifetime, it is optional. It can be identified by name, but also by index (first 'reg' is CCSR). Note that while all SoCs should have a DCFG register block in DCSR, we only need to expose it for the SoCs where the RCW override procedure is known to be needed and has been validated. Signed-off-by: Vladimir Oltean <vladimir.oltean@nxp.com> Reviewed-by: Conor Dooley <conor.dooley@microchip.com> Link: https://lore.kernel.org/r/20260721231603.67865-8-vladimir.oltean@nxp.com Signed-off-by: Christophe Leroy (CS GROUP) <chleroy@kernel.org>
2026-08-07soc: fsl: guts: make fsl_soc_data available after fsl_guts_init()Vladimir Oltean1-7/+10
In a future change, struct fsl_soc_data will be extended with methods for performing RCW override. Since this will be performed from a calling context outside fsl_guts_init(), we need to keep track of the soc_data that we determine at fsl_guts_init() time, so we can reference it later. Signed-off-by: Vladimir Oltean <vladimir.oltean@nxp.com> Reviewed-by: Ioana Ciornei <ioana.ciornei@nxp.com> Link: https://lore.kernel.org/r/20260721231603.67865-7-vladimir.oltean@nxp.com Signed-off-by: Christophe Leroy (CS GROUP) <chleroy@kernel.org>
2026-08-07soc: fsl: guts: make it easier to determine on which SoC we are runningIoana Ciornei1-6/+41
The guts driver will need to easily determine on which SoC it's running when it will need to perform RCW override at runtime. The guts driver knows this already because fsl_guts_init() reads the QorIQ/Layerscape architectural System Version Register (SVR), but it doesn't save this for later lookups. Add a new qoriq_die enum to be used as an index in the fsl_soc_die array. A new fsl_soc_die_match_one() function is also added so that we can directly determine if the SVR is a match with a specific die. The SVR value read from the DCFG CCSR is also kept in the global soc structure so that it can be accessed when needed. Signed-off-by: Ioana Ciornei <ioana.ciornei@nxp.com> Signed-off-by: Vladimir Oltean <vladimir.oltean@nxp.com> Link: https://lore.kernel.org/r/20260721231603.67865-6-vladimir.oltean@nxp.com Signed-off-by: Christophe Leroy (CS GROUP) <chleroy@kernel.org>
2026-08-07soc: fsl: guts: add a central fsl_guts_read() functionIoana Ciornei1-4/+9
Add a central fsl_guts_read() function which will take into account the endianness that was already determined. No point is duplicating the if-else statement each time we need to read a DCFG register. Signed-off-by: Ioana Ciornei <ioana.ciornei@nxp.com> Signed-off-by: Vladimir Oltean <vladimir.oltean@nxp.com> Link: https://lore.kernel.org/r/20260721231603.67865-5-vladimir.oltean@nxp.com Signed-off-by: Christophe Leroy (CS GROUP) <chleroy@kernel.org>
2026-08-07soc: fsl: guts: add a global structure to hold stateIoana Ciornei1-11/+18
Add the fsl_soc_guts structure in order to pass information like base addresses, endianness etc between the init time and the runtime operations (RCW override) which will get added in future patches. There is no point in mapping and unmapping the DCFG CCSR space every time we need to make a read, just map it once and keep its reference in this new global structure. Signed-off-by: Ioana Ciornei <ioana.ciornei@nxp.com> Signed-off-by: Vladimir Oltean <vladimir.oltean@nxp.com> Link: https://lore.kernel.org/r/20260721231603.67865-4-vladimir.oltean@nxp.com [chleroy: fixed typo on 'structure' in commit message] Signed-off-by: Christophe Leroy (CS GROUP) <chleroy@kernel.org>
2026-08-07soc: fsl: guts: use a macro to encode the DCFG CCSR spaceIoana Ciornei1-1/+3
Instead of using a hardcoded value when iomapping the DCFG CCSR space, add a new macro for it. The code will be easier to follow this way, especially when we add support for the DCFG DCSR space as well. Signed-off-by: Ioana Ciornei <ioana.ciornei@nxp.com> Signed-off-by: Vladimir Oltean <vladimir.oltean@nxp.com> Link: https://lore.kernel.org/r/20260721231603.67865-3-vladimir.oltean@nxp.com Signed-off-by: Christophe Leroy (CS GROUP) <chleroy@kernel.org>
2026-08-07soc: fsl: guts: perform fsl_guts_init() error teardown in reverse order of setupVladimir Oltean1-13/+20
fsl_guts_init() is about to get much more complicated and the central error handling procedure cannot scale in its current design, unless we add a lot of "if" conditions to detect what has been allocated and what hasn't. Currently the code relies on the fact that kfree(NULL) is safe, but this doesn't scale to the case where "soc_dev_attr" itself is NULL, because this would dereference "soc_dev_attr->family" and friends of a NULL pointer. Convert to the more typical error handling pattern where the teardown is in the strict reverse order of setup, and a teardown step is only called if its corresponding setup step was executed. At the same time, maintain the optionality of soc_dev_attr->serial_number by not checking whether that kasprintf() has returned NULL. In the error path, kfree(NULL) is safe, so we don't need to add an "if" condition for it. Michael Walle has confirmed that ignoring the error was intentional, and we preserve that: https://lore.kernel.org/linux-phy/DK44809N7Y8I.J2Z3U4N32H0Q@kernel.org/ Signed-off-by: Vladimir Oltean <vladimir.oltean@nxp.com> Reviewed-by: Ioana Ciornei <ioana.ciornei@nxp.com> Link: https://lore.kernel.org/r/20260721231603.67865-2-vladimir.oltean@nxp.com Signed-off-by: Christophe Leroy (CS GROUP) <chleroy@kernel.org>
2026-08-07soc: fsl: dpio: fix kernel-doc typosRandy Dunlap1-2/+2
Correct spelling of 2 words. Signed-off-by: Randy Dunlap <rdunlap@infradead.org> Cc: Li Yang <leoyang.li@nxp.com> Cc: linuxppc-dev@lists.ozlabs.org Cc: linux-arm-kernel@lists.infradead.org Cc: Frank Li <Frank.Li@nxp.com> Cc: Guanhua Gao <guanhua.gao@nxp.com> Cc: Roy Pledge <Roy.Pledge@nxp.com> Reviewed-by: Ioana Ciornei <ioana.ciornei@nxp.com> Link: https://lore.kernel.org/r/20260728004938.905415-1-rdunlap@infradead.org Signed-off-by: Christophe Leroy (CS GROUP) <chleroy@kernel.org>
2026-08-07soc: fsl: fix kernel-doc warnings and typosRandy Dunlap1-9/+10
Correct spelling of "list". Fix a kernel-doc warning by describing the nested structure completely: include/soc/fsl/dpaa2-fd.h:52: warning: Function parameter or member 'simple' not described in 'dpaa2_fd' Signed-off-by: Randy Dunlap <rdunlap@infradead.org> Cc: Li Yang <leoyang.li@nxp.com> Cc: linuxppc-dev@lists.ozlabs.org Cc: linux-arm-kernel@lists.infradead.org Cc: Frank Li <Frank.Li@nxp.com> Cc: Guanhua Gao <guanhua.gao@nxp.com> Cc: Roy Pledge <Roy.Pledge@nxp.com> Reviewed-by: Ioana Ciornei <ioana.ciornei@nxp.com> Link: https://lore.kernel.org/r/20260728004924.904210-1-rdunlap@infradead.org Signed-off-by: Christophe Leroy (CS GROUP) <chleroy@kernel.org>
2026-08-07bus: fsl-mc: Remove redundant dev_err()Pan Chuang1-5/+1
Since commit 55b48e23f5c4 ("genirq/devres: Add error handling in devm_request_*_irq()"), devm_request_threaded_irq() automatically logs detailed error messages on failure. Remove the now-redundant driver-specific dev_err() calls. Signed-off-by: Pan Chuang <panchuang@vivo.com> Reviewed-by: Ioana Ciornei <ioana.ciornei@nxp.com> Link: https://lore.kernel.org/r/20260710110930.462109-2-panchuang@vivo.com Signed-off-by: Christophe Leroy (CS GROUP) <chleroy@kernel.org>
2026-08-07soc: fsl: qe: Add support of IRQs in QE GPIOPaul Louvel2-1/+28
Some QE GPIO pins have an associated interrupt line in the QE PIC to signal state changes on the pin. Because the GPIO controller does not perform any interrupt handling itself, a nexus node (interrupt-map) is used to map each GPIO line supporting IRQ to the parent QE PIC interrupt domain. Add the to_irq() method in the corresponding GPIO controller driver, that uses the nexus node to perform the translation. Signed-off-by: Paul Louvel <paul.louvel@bootlin.com> Reviewed-by: Christophe Leroy (CS GROUP) <chleroy@kernel.org> Link: https://lore.kernel.org/r/20260708-qe-pic-gpios-v2-10-1972044cfbd1@bootlin.com [chleroy: Added dependency on !SPARC to fix build on sparc reported by 0-day kernel robot] Signed-off-by: Christophe Leroy (CS GROUP) <chleroy@kernel.org>
2026-08-06Merge tag 'mm-hotfixes-stable-2026-08-06-18-44' of ↵Linus Torvalds19-165/+246
git://git.kernel.org/pub/scm/linux/kernel/git/akpm/mm Pull MM fixes from Andrew Morton: "17 hotfixes. 15 are cc:stable. 16 are for MM. There's a patch series from Lorenzo "mm: fix UAF caused by race between ptdump and vmap pgtable freeing" which addresses a quite old bug in the ptdump code. And another series also from Lorenzo which fixes a four year old bug in the huge_zero_folio handling. A series from SJ fixes a few possible divide-by-zero issues which Sashiko sniffed out. And a series which fixes handling of the commit_inputs parameters. The remainder are singletons, please see their changelogs for details" * tag 'mm-hotfixes-stable-2026-08-06-18-44' of git://git.kernel.org/pub/scm/linux/kernel/git/akpm/mm: mm/damon: adjust isolated pages stat for DAMOS_MIGRATE_{HOT,COLD} mm/damon/ops-common: putback folios on invalid migrate nid mm/huge_memory: initialise workingset state before folio split mm/page_table_check: skip special zero mappings mm/damon/lru_sort: skip damon_call() if ctx has not started mm/damon/reclaim: skip damon_call() if ctx has not started mm/damon/lru_sort: error out for >10000 active_mem_bp samples/damon/mtier: error out for zero quota goal target values mailmap: map old addresses to Danila Tikhonov mm/huge_memory: separate out CONFIG_PERSISTENT_HUGE_ZERO_FOLIO logic mm/huge_memory: fix huge_zero_pfn race MAINTAINERS: update address for Brendan Jackman mm/filemap: __filemap_add_folio() restore index before retrying microblaze: restore the page alignment of swapper_pg_dir arm64: remove redundant concurrent ptdump UAF mitigation mm/ptdump: always stabilise against page table freeing using init_mm mm/vmalloc: acquire init_mm lock on huge vmap to avoid ptdump UAF
2026-08-06Merge tag 'v7.2-rc6-smb3-server-fixes' of git://git.samba.org/ksmbdLinus Torvalds5-11/+32
Pull smb server fixes from Steve French: - Reject Pattern_V1 payloads when Pattern_V1 support was not negotiated - Validate compression transform flags and chained mode before allocating the decompression buffer - Enforce the pre-authentication PDU size limit before allocating the decompression buffer, preventing compressed requests from bypassing the limit * tag 'v7.2-rc6-smb3-server-fixes' of git://git.samba.org/ksmbd: ksmbd: apply the pre-authentication PDU limit when decompressing ksmbd: validate compression Flags before kvmalloc smb: compress: reject Pattern_V1 when not negotiated
2026-08-06riscv: ftrace: Fix ftrace_modify_call failure on kprobed functionsPu Lehui1-1/+4
We are frequently hitting the following splat during the riscv bpf selftests: 00000000026dc75a: expected (7c3ff297) but got (00100073) ------------[ ftrace bug ]------------ ftrace failed to modify [<ffffffff03c44c1c>] bpf_kfunc_common_test+0x4/0x20 [bpf_testmod] actual: e7:82:c2:ce Updating ftrace call site to call a different ftrace function ftrace record flags: 80100002 (2) expected tramp: ffffffff80043904 ------------[ cut here ]------------ WARNING: kernel/trace/ftrace.c:2278 at ftrace_bug+0x46e/0x4b0, CPU#1: test_progs/98 ... [<ffffffff80008f4e>] ftrace_bug+0x46e/0x4b0 [<ffffffff803d3e86>] ftrace_replace_code+0x16e/0x170 [<ffffffff803d42b6>] ftrace_modify_all_code+0x12e/0x1b8 [<ffffffff800430f4>] arch_ftrace_update_code+0x14/0x28 [<ffffffff803e0324>] ftrace_startup+0x14c/0x2a0 [<ffffffff803e133c>] ftrace_startup_subops+0x584/0x1050 [<ffffffff804500e6>] register_ftrace_graph+0x4e6/0x1018 [<ffffffff804cf9f6>] register_fprobe_ips+0xc66/0x12f8 [<ffffffff8049abe8>] bpf_kprobe_multi_link_attach+0x5d8/0xe68 [<ffffffff8050fcaa>] __sys_bpf+0x3d5a/0x47f0 [<ffffffff805107ee>] __riscv_sys_bpf+0xae/0x168 [<ffffffff80034d78>] syscall_handler+0x60/0x100 [<ffffffff8228b4f4>] do_trap_ecall_u+0x174/0x208 [<ffffffff822b69c4>] handle_exception+0x16c/0x178 After debugging, it can be triggered by similar commands below: ``` echo do_nanosleep > set_ftrace_filter echo function > current_tracer echo 'p do_nanosleep' > kprobe_events echo 1 > events/kprobes/enable echo 'f do_nanosleep' > dynamic_events echo 1 > events/fprobes/enable ``` The reason is that attaching a kprobe to an ftrace-traced function entry replaces its initial auipc insn with ebreak. When ftrace_modify_call later runs, it expects auipc insn, so verification fails and triggers ftrace_bug. The expected auipc logic remains conceptually unchanged, and kprobe single-stepping ensures normal execution. Therefore, if the first insn is ebreak, bypassing the check to continue patching the jalr insn is safe and avoids ftrace failures. Fixes: b2137c3b6d7a ("riscv: ftrace: prepare ftrace for atomic code patching") Signed-off-by: Pu Lehui <pulehui@huawei.com> Link: https://patch.msgid.link/20260802094929.3978390-1-pulehui@huaweicloud.com [pjw@kernel.org: fixed reproducer in commit message] Signed-off-by: Paul Walmsley <pjw@kernel.org>
2026-08-06ext4: fix estimate extent index blocks in ext4_ext_index_trans_blocks()Jan Kara1-3/+11
The estimate of the number of impacted extent tree index blocks could be one-too-low. If we modify say 2 extents, already two leaf index blocks could be impacted, not just one the current estimate counts with. Fix the estimate. Signed-off-by: Jan Kara <jack@suse.cz> Link: https://patch.msgid.link/20260805153605.166545-6-jack@suse.cz Signed-off-by: Theodore Ts'o <tytso@mit.edu>
2026-08-06ext4: fix transaction overflow during writebackJan Kara1-3/+3
Commit 95ad8ee45cdb ("ext4: correct the reserved credits for extent conversion") was correct to note that we need to reserve enough credits for all extents possibly underlying a large folio. However it was too eager to reduce the number of reserved credits. Extent conversion may not only need to touch several leaf extent blocks, it may also need to split extents - for example a single large unwritten extent may need to be split into many small written ones in case of sparse folio dirtying. This can thus result not only in extent leaf modifications but also in a need to allocate new extent tree nodes. As a result the reserved transaction credits were not sufficient in some corner cases. Use ext4_meta_trans_blocks() for correct upper bound credit estimate. Fixes: 95ad8ee45cdb ("ext4: correct the reserved credits for extent conversion") Signed-off-by: Jan Kara <jack@suse.cz> Link: https://patch.msgid.link/20260805153605.166545-5-jack@suse.cz Signed-off-by: Theodore Ts'o <tytso@mit.edu>
2026-08-06ext4: teach ext4_meta_trans_blocks() about number of allocated extentsJan Kara3-17/+17
So far ext4_meta_trans_blocks() expects that each extent counted in @pextents will be allocated in the transaction we estimate credits for. This is correct for the use in ext4_chunk_trans_blocks() and ext4_chunk_trans_extent() however the use in atomic write path (ext4_convert_unwritten_extents_atomic() and ext4_iomap_alloc() for IOMAP_ATOMIC) unnecessarily overestimates the number of necessary credits as neither of them allocates any data. Add argument to ext4_meta_trans_blocks() for number of extents that are going to be allocated in the transaction. Signed-off-by: Jan Kara <jack@suse.cz> Link: https://patch.msgid.link/20260805153605.166545-4-jack@suse.cz Signed-off-by: Theodore Ts'o <tytso@mit.edu>
2026-08-06selftests/mm: unpoison pages in memory-failure teardownMuhammad Usama Anjum1-22/+22
The memory-failure tests call cleanup() only after all result checks. A failed ASSERT_* invokes fixture teardown and aborts the test, so it skips cleanup() and leaves the injected page hardware-poisoned. Invoke cleanup() from FIXTURE_TEARDOWN() instead. Guard it with self->injection_attempted so tests that exit before injection do not try to unpoison a page when no injection was attempted. Injection can poison a page before returning an error or delivering SIGBUS, so teardown must clean up after every injection attempt. This runs the existing HWPoison and HardwareCorrupted checks on both normal and assertion-failure paths. Link: https://lore.kernel.org/20260729091127.1001179-1-usama.anjum@arm.com Fixes: ff4ef2fbd101 ("selftests/mm: add memory failure anonymous page test") Signed-off-by: Muhammad Usama Anjum <usama.anjum@arm.com> Reviewed-by: David Hildenbrand (Arm) <david@kernel.org> Acked-by: Miaohe Lin <linmiaohe@huawei.com> Cc: Liam R. Howlett <liam@infradead.org> Cc: Lorenzo Stoakes <ljs@kernel.org> Cc: Michal Hocko <mhocko@suse.com> Cc: Mike Rapoport <rppt@kernel.org> Cc: Naoya Horiguchi <nao.horiguchi@gmail.com> Cc: Shuah Khan <shuah@kernel.org> Cc: Suren Baghdasaryan <surenb@google.com> Cc: Vlastimil Babka <vbabka@kernel.org> Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
2026-08-06mm/shmem: downgrade final i_blocks check in shmem_evict_inode() to pr_warn()Jiacheng Yu1-1/+4
shmem_evict_inode() ends with WARN_ON(inode->i_blocks) as a final consistency check of shmem's block accounting. When it fires, the inode-local counters die with the inode; what may linger is a small residue in accounting kept outside the inode, such as per-mount or per-user charges. No data is lost, and no corruption follows. On kernels running with panic_on_warn=1, this accounting inconsistency escalates to a full machine panic, which is disproportionate to the impact. Downgrade the WARN_ON() to a pr_warn() that reports the inode together with its accounting counters (i_blocks, alloced, swapped, nrpages), keeping the inconsistency visible in the logs. The accounting bugs this check has caught over the years -- the swapout race described in commit 0f3c42f522dc ("tmpfs: change final i_blocks BUG to WARNING") and the error recovery race fixed in commit 267a4c76bbdb ("tmpfs: fix shmem_evict_inode() warnings on i_blocks") -- are real and should still be fixed; this change only removes the disproportionate escalation. One way to hit this race: soft_offline_in_use_page()'s fast path drops a clean, unmapped shmem folio via mapping_evict_folio(), where the xas_store() and the nrpages decrement are not atomic against a concurrent shmem_evict_inode(); the final shmem_recalc_inode() can then read the pre-decrement nrpages, compute freed = 0, and leave one page charged. Same class as the races in 0f3c42f522dc and 267a4c76bbdb, this time in the under-count direction; reproduced on 7.2-rc4 with madvise(MADV_SOFT_OFFLINE) racing MAP_FIXED replacement of a shared-anonymous VMA. [yujiacheng3@huawei.com: drop redundant casts in shmem_evict_inode() pr_warn] Link: https://lore.kernel.org/20260729121201.776566-1-yujiacheng3@huawei.com Link: https://lore.kernel.org/20260728091014.3876715-1-yujiacheng3@huawei.com Fixes: 0f3c42f522dc ("tmpfs: change final i_blocks BUG to WARNING") Signed-off-by: Jiacheng Yu <yujiacheng3@huawei.com> Cc: Baolin Wang <baolin.wang@linux.alibaba.com> Cc: Hugh Dickins <hughd@google.com> Cc: Yongqiang Liu <liuyongqiang13@huawei.com> Cc: Christian Brauner <brauner@kernel.org> Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
2026-08-06mm/khugepaged: replace mutex_lock/mutex_unlock usage with guard macroJakov Novak1-16/+15
Currently, khugepaged locks the khugepaged_mutex in two functions: start_stop_khugepaged and khugepaged_min_free_kbytes_update. Remove mutex_lock/mutex_unlock usage in these functions and replace it with the guard macro. This makes the code more readable (removing a goto statement) and makes it harder to introduce bugs in the future. No functional changes introduced. Link: https://lore.kernel.org/20260730204724.16912-1-jakovnovak30@gmail.com Signed-off-by: Jakov Novak <jakovnovak30@gmail.com> Reviewed-by: Dev Jain <dev.jain@arm.com> Reviewed-by: Lorenzo Stoakes (ARM) <ljs@kernel.org> Reviewed-by: Andrew Morton <akpm@linux-foundation.org> Reviewed-by: Zi Yan <ziy@nvidia.com> Acked-by: David Hildenbrand (Arm) <david@kernel.org> Cc: Baolin Wang <baolin.wang@linux.alibaba.com> Cc: Barry Song <baohua@kernel.org> Cc: Lance Yang <lance.yang@linux.dev> Cc: Liam R. Howlett <liam@infradead.org> Cc: Nico Pache <npache@redhat.com> Cc: Ryan Roberts <ryan.roberts@arm.com> Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
2026-08-06mm/zsmalloc: fix release order of locks in zs_page_migrate()Richard Chang1-2/+2
In zs_page_migrate(), locks are acquired in the following order: 1. write_lock(&pool->lock) 2. spin_lock(&class->lock) 3. zspage_write_trylock(zspage) However, upon successful page migration, they were being released in forward acquisition (FIFO) order: 1. write_unlock(&pool->lock) 2. spin_unlock(&class->lock) 3. zspage_write_unlock(zspage) Fix the unlocking order to release locks in strict reverse (LIFO) order of acquisition: 3. zspage_write_unlock(zspage) 2. spin_unlock(&class->lock) 1. write_unlock(&pool->lock) Releasing locks in reverse order of acquisition adheres to standard kernel locking hygiene, prevents potential lock ordering and lockdep inconsistencies. Link: https://lore.kernel.org/20260728055333.421080-1-richardycc@google.com Signed-off-by: Richard Chang <richardycc@google.com> Reviewed-by: Sergey Senozhatsky <senozhatsky@chromium.org> Tested-by: Sergey Senozhatsky <senozhatsky@chromium.org> Cc: Martin Liu <liumartin@google.com> Cc: Minchan Kim <minchan@kernel.org> Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
2026-08-06Documentation: zram: remove sections numberingSergey Senozhatsky1-20/+20
Those numbers are difficult to maintain and in fact we can refer to sections by their names (in html). Link: https://lore.kernel.org/20260728021229.181627-1-senozhatsky@chromium.org Signed-off-by: Sergey Senozhatsky <senozhatsky@chromium.org> Suggested-by: Sebastian Andrzej Siewior <bigeasy@linutronix.de> Reviewed-by: SJ Park <sj@kernel.org> Cc: Jonathan Corbet <corbet@lwn.net> Cc: Minchan Kim <minchan@kernel.org> Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
2026-08-06ksm: stop iterating VMAs when ksm_test_exit returns trueWang Wensheng1-1/+1
In scan_get_next_rmap_item() the break statement only exits the inner while loop, leaving remaining VMAs to be iterated even if ksm_test_exit() returns true. Replace it with a goto statement to avoid the unnecessary work. Link: https://lore.kernel.org/20260726133501.504048-1-wsw9603@163.com Signed-off-by: Wang Wensheng <wsw9603@163.com> Reviewed-by: Xu Xin <xu.xin16@zte.com.cn> Acked-by: David Hildenbrand (Arm) <david@kernel.org> Cc: Chengming Zhou <chengming.zhou@linux.dev> Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
2026-08-06mm: fold userfaultfd_rwp() to false without CONFIG_ARCH_HAS_PTE_PROTNONEKiryl Shutsemau (Meta)1-0/+6
RWP tracks accesses by installing PAGE_NONE (protnone) PTEs, so its code paths are gated on userfaultfd_rwp(). Without CONFIG_ARCH_HAS_PTE_PROTNONE there is no PAGE_NONE -- <linux/pgtable.h> defines it to a BUILD_BUG() stub, relying on callers folding such paths to dead code via IS_ENABLED(CONFIG_ARCH_HAS_PTE_PROTNONE). userfaultfd_rwp() was not a compile-time constant, so the compiler could not fold those paths. With an older compiler (gcc 8.5.0, sparc64) the PAGE_NONE reference in move_pages_huge_pmd() survived to codegen: mm/huge_memory.c:2874: _dst_pmd = pmd_modify(_dst_pmd, PAGE_NONE); compiler_types.h:702: error: call to '__compiletime_assert_501' declared with attribute error: BUILD_BUG failed RWP cannot exist without protnone, so return a compile-time false when CONFIG_ARCH_HAS_PTE_PROTNONE is unset; every RWP path then folds away. Link: https://lore.kernel.org/amcitKvUvFYr8W38@thinkstation Signed-off-by: Kiryl Shutsemau (Meta) <kas@kernel.org> Reported-by: kernel test robot <lkp@intel.com> Closes: https://lore.kernel.org/oe-kbuild-all/202607250853.VaJWGLeA-lkp@intel.com/ Cc: Andrea Arcangeli <aarcange@redhat.com> Cc: David Hildenbrand <david@kernel.org> Cc: James Houghton <jthoughton@google.com> Cc: Jonathan Corbet <corbet@lwn.net> Cc: Liam Howlett <liam@infradead.org> Cc: Lorenzo Stoakes <ljs@kernel.org> Cc: Mike Rapoport (Microsoft) <rppt@kernel.org> Cc: Paolo Bonzini <pbonzini@redhat.com> Cc: Peter Xu <peterx@redhat.com> Cc: Sean Christopherson <seanjc@google.com> Cc: SeongJae Park <sj@kernel.org> Cc: Suren Baghdasaryan <surenb@google.com> Cc: Vlastimil Babka <vbabka@kernel.org> Cc: Zi Yan <ziy@nvidia.com> Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
2026-08-06mm/migrate: report RCU-tasks quiescent states in migrate_pages_batch()Breno Leitao1-1/+1
migrate_pages_batch() unmaps each folio before moving it, and every unmap runs the mmu_notifier invalidate callbacks. On KVM hosts try_to_migrate() ends up in kvm_mmu_notifier_invalidate_range_start() -> tdp_mmu_zap_leafs(), which is expensive, so unmapping a large batch keeps the CPU busy for a long time. The loop already calls cond_resched(), but on PREEMPTION kernels that is a no-op, and involuntary preemption is not a Tasks-RCU quiescent state. A long batch therefore never reports a quiescent state, and the migrating task (e.g. kcompactd) becomes a Tasks-RCU holdout, stalling the Tasks-RCU grace period for minutes, which is common at Meta fleet: INFO: rcu_tasks detected stalls on tasks: 0000000055349ecc: .. nvcsw: 1157401/1157401 holdout: 1 idle_cpu: -1/56 task:kcompactd0 state:R running task Call Trace: tdp_mmu_zap_leafs tdp_mmu_next_root gfn_to_pfn_cache_invalidate_start kvm_mmu_notifier_invalidate_range_start __mmu_notifier_invalidate_range_start try_to_migrate_one try_to_migrate migrate_pages_batch migrate_pages compact_zone compact_node kcompactd kthread Use cond_resched_tasks_rcu_qs() so a quiescent state is reported even when cond_resched() does nothing. This has also been discussed at [1] Link: https://lore.kernel.org/20260727-kcompact-v1-1-bdfefddd6874@debian.org Link: https://lore.kernel.org/all/amdWVTs0WKOxguxP@gmail.com/ [1] Signed-off-by: Breno Leitao <leitao@debian.org> Acked-by: Zi Yan <ziy@nvidia.com> Reviewed-by: Gregory Price <gourry@gourry.net> Reviewed-by: Paul E. McKenney <paulmck@kernel.org> Acked-by: David Hildenbrand (Arm) <david@kernel.org> Cc: Alistair Popple <apopple@nvidia.com> Cc: Byungchul Park <byungchul@sk.com> Cc: "Huang, Ying" <ying.huang@linux.alibaba.com> Cc: Joshua Hahn <joshua.hahnjy@gmail.com> Cc: Matthew Brost <matthew.brost@intel.com> Cc: Rakie Kim <rakie.kim@sk.com> Cc: <stable@vger.kernel.org> Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
2026-08-06zram: use a custom key for each zram objectSebastian Andrzej Siewior2-8/+4
Each struct zram uses the same key for its struct lockdep_map which is used for locking analysis. According to Sergey the lock chains might be different if zram1 is used for and zram2 is for ext4. This might lead to false dead lock reports if it mixes a zram1 chain with a zram2. This can be avoided if each lockmap gets its own unique key.c Use a dynamic lock_class_key for the table_lock_map. Link: https://lore.kernel.org/20260714141300.3945672-3-bigeasy@linutronix.de Signed-off-by: Sebastian Andrzej Siewior <bigeasy@linutronix.de> Reviewed-by: Sergey Senozhatsky <senozhatsky@chromium.org> Tested-by: Sergey Senozhatsky <senozhatsky@chromium.org> Cc: Jens Axboe <axboe@kernel.dk> Cc: Minchan Kim <minchan@kernel.org> Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
2026-08-06zram: move lockmap to be per-zram instead per tableSebastian Andrzej Siewior2-13/+10
Patch series "zram: lockmap tweaks". This patch (of 2): The zram object contains an array zram_table_entry. Each one has a `lock' variable and each has a matching struct lockdep_map. This mimics a struct mutex. It uses always the same key for all lockdep_map instances. This makes it look like the same lock to lockdep. Therefore it could be reduced to have one lockdep_map per struct zram. Use only one struct lockdep_map per struct zram. Link: https://lore.kernel.org/20260714141300.3945672-1-bigeasy@linutronix.de Link: https://lore.kernel.org/20260714141300.3945672-2-bigeasy@linutronix.de Signed-off-by: Sebastian Andrzej Siewior <bigeasy@linutronix.de> Reviewed-by: Sergey Senozhatsky <senozhatsky@chromium.org> Tested-by: Sergey Senozhatsky <senozhatsky@chromium.org> Cc: Jens Axboe <axboe@kernel.dk> Cc: Minchan Kim <minchan@kernel.org> Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
2026-08-06selftests/mm: fix gup_longterm EINVAL error messagezhaozhengzhuo1-1/+1
The gup_longterm test prints a literal "n" when PIN_LONGTERM_TEST_START fails with EINVAL because the string is missing the newline escape sequence. Print a newline instead. Link: https://lore.kernel.org/23557F4CB8CF36FF+20260724074603.1479243-1-zhaozhengzhuo@uniontech.com Signed-off-by: zhaozhengzhuo <zhaozhengzhuo@uniontech.com> Reviewed-by: Sarthak Sharma <sarthak.sharma@arm.com> Reviewed-by: Dev Jain <dev.jain@arm.com> Acked-by: David Hildenbrand (arm) <david@kernel.org> Cc: Jason Gunthorpe <jgg@ziepe.ca> Cc: John Hubbard <jhubbard@nvidia.com> Cc: Peter Xu <peterx@redhat.com> Cc: Shuah Khan <shuah@kernel.org> Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
2026-08-06mm: page_alloc: fix non-movable reclaim storm in defrag_modeJohannes Weiner3-6/+33
As we deployed defrag_mode into Meta production, pressure spikes and excessive swapping were observed on some workloads. Tracing confirmed that this is unmovable/reclaimable requests spinning in the allocator and direct reclaim, causing excessive amounts of swap. The initial plan for defrag_mode was to rely on kswapd/kcompactd to produce blocks, and if those are overwhelmed under high pressure, let the allocator fall back (__rmqueue_steal()) after its retry loops. However, that retrying results in more reclaim on some of these workloads than we'd hoped, sometimes excessively so, spurred on by the !costly order conditions in should_reclaim_retry(). The storms are dependent on the request type. Reclaim will inevitably make room in existing movable blocks, since that's where the LRU pages live. So if movable requests retry on reclaim, they make progress. When non-movable requests spin in reclaim that isn't productive. They cannot use the individually freed pages, and the process is unlikely to accidentally free whole blocks to meet the ALLOC_NOFRAGMENT bar. They spin and overreclaim excessively, which tanks performance and triggers userspace guards like swap exhaustion or pressure based OOM. To fix this, send non-movable requests, regardless of order, into pageblock reclaim/compaction. This way, they help move things along to meet the ALLOC_NOFRAGMENT bar. After this patch, the reclaim storms and excess OOM rates are no longer observed in production. The longer-term plan is still to have all requests, including the movable ones, help make blocks to spread the cost of defragmenting more evenly and fairly; combined with proper watermarking to reduce allocation latencies in the common case. However, doing this naively unearths scaling and concurrency limitations in compaction that need to be addressed first. Promoting just non-movables for now is the minimally viable bug fix for the above issue. [brendan.jackman@linux.dev: fix try_to_compact_pages() kerneldoc] Link: https://lore.kernel.org/DK7NM9RPUJOD.11PNJJ5N2OBED@linux.dev Link: https://lore.kernel.org/20260722150006.3848560-5-hannes@cmpxchg.org Fixes: e3aa7df331bc ("mm: page_alloc: defrag_mode") Signed-off-by: Johannes Weiner <hannes@cmpxchg.org> Signed-off-by: "Brendan Jackman" <brendan.jackman@linux.dev> Reviewed-by: Vlastimil Babka (SUSE) <vbabka@kernel.org> Cc: Brendan Jackman <brendan.jackman@linux.dev> Cc: David Hildenbrand <david@kernel.org> Cc: Gregory Price <gourry@gourry.net> Cc: Liam R. Howlett <liam@infradead.org> Cc: Lorenzo Stoakes <ljs@kernel.org> Cc: Michal Hocko <mhocko@suse.com> Cc: Mike Rapoport <rppt@kernel.org> Cc: Shakeel Butt <shakeel.butt@linux.dev> Cc: Suren Baghdasaryan <surenb@google.com> Cc: Zi Yan <ziy@nvidia.com> Cc: <stable@vger.kernel.org> Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
2026-08-06mm: page_alloc: move capture_control to the page allocatorVlastimil Babka (SUSE)4-45/+57
The compaction capturing code assumes the allocation request order and compaction target order are the same. That won't be true once defrag_mode promotes sub-block allocations to pageblock-order compaction: compaction targets the larger order, while capture should remain at the original allocation order. Move the capture_control to the page allocator and give it its own copies of what the page freeing path matches against - zone, migratetype and the allocation order - rather than reaching into compaction's live compact_control. __alloc_pages_direct_compact() fills in migratetype and order, and installs and hides current->capture_control around the whole compaction call; try_to_compact_pages() aims capc->zone at each zone while it is being compacted. compact_zone_order() no longer deals with capture at all. Pass the capture_control through try_to_compact_pages() / compact_zone_order() in place of the bare struct page **. No functional change. Link: https://lore.kernel.org/20260722150006.3848560-4-hannes@cmpxchg.org Fixes: e3aa7df331bc ("mm: page_alloc: defrag_mode") Signed-off-by: Vlastimil Babka (SUSE) <vbabka@kernel.org> Co-developed-by: Johannes Weiner <hannes@cmpxchg.org> Signed-off-by: Johannes Weiner <hannes@cmpxchg.org> Reviewed-by: Gregory Price <gourry@gourry.net> Cc: Brendan Jackman <brendan.jackman@linux.dev> Cc: Brendan Jackman <jackmanb@google.com> Cc: David Hildenbrand <david@kernel.org> Cc: Liam R. Howlett <liam@infradead.org> Cc: Lorenzo Stoakes <ljs@kernel.org> Cc: Michal Hocko <mhocko@suse.com> Cc: Mike Rapoport <rppt@kernel.org> Cc: Shakeel Butt <shakeel.butt@linux.dev> Cc: Suren Baghdasaryan <surenb@google.com> Cc: Zi Yan <ziy@nvidia.com> Cc: <stable@vger.kernel.org> Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
2026-08-06mm: compaction: support non-movable compaction for pageblock requestsJohannes Weiner1-7/+39
While trying to fix a reclaim storm in defrag_mode, I noticed that non-movable direct compaction is extremely inefficient. When searching for space to evacuate, compaction only allows blocks of the same type as the incoming request. This is to prevent migratetype pollution, where a small non-movable request frees space in a movable block and provokes the allocator to fall back and pollute it. This protection is reasonable on one hand, but the downside is that it makes non-movable direct compaction nearly useless: if we get the type annotations right, by definition there aren't any movable pages inside the non-movable blocks it is allowed to scan. With defrag_mode, the goal is the production of whole blocks, which are essentially type neutral: __rmqueue_claim() will convert them wholesale on alloc. This makes type mixing and pollution a non-issue. Fix the pollution gates to take the requested order into account, and allow whole-block requests to scan blocks of other types. The only exception is CMA blocks. That type is sticky and these blocks cannot be claimed to other types. Continue to be strict with them, and allow only explicit ALLOC_CMA requests and kcompactd to evacuate them. Link: https://lore.kernel.org/20260722150006.3848560-3-hannes@cmpxchg.org Fixes: e3aa7df331bc ("mm: page_alloc: defrag_mode") Signed-off-by: Johannes Weiner <hannes@cmpxchg.org> Reviewed-by: Vlastimil Babka (SUSE) <vbabka@kernel.org> Reviewed-by: Gregory Price <gourry@gourry.net> Cc: Brendan Jackman <brendan.jackman@linux.dev> Cc: Brendan Jackman <jackmanb@google.com> Cc: David Hildenbrand <david@kernel.org> Cc: Liam R. Howlett <liam@infradead.org> Cc: Lorenzo Stoakes <ljs@kernel.org> Cc: Michal Hocko <mhocko@suse.com> Cc: Mike Rapoport <rppt@kernel.org> Cc: Shakeel Butt <shakeel.butt@linux.dev> Cc: Suren Baghdasaryan <surenb@google.com> Cc: Zi Yan <ziy@nvidia.com> Cc: <stable@vger.kernel.org> Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
2026-08-06mm: page_alloc: __GFP_FS lockdep annotation for direct compactionJohannes Weiner1-0/+2
Patch series "mm: fix reclaim storms in defrag_mode", v2. As we deployed vm.defrag_mode=1 in Meta production, some workloads regressed with recurring pressure spikes and swap storms (which in turn triggered userspace OOM rules on pressure and swap utilization levels). Tracing pinned this to non-movable requests spinning and reclaiming unproductively when kswapd/kcompactd are overwhelmed. Direct reclaim predominantly frees up pages in movable blocks, but those requests cannot use that space under defrag_mode rules; and it is unlikely to free up whole blocks incidentally for __rmqueue_claim() to work. This series fixes it by making non-movable requests participate in pageblock production in the allocator slowpath - meaning, they will invoke direct reclaim and direct compaction with pageblock_order. That requires some small-ish adjustments up front in the allocator and the compaction code: three prep patches and the fix last. The series has been in production against one of the affected workloads for several weeks and restores the OOM kill rate to !defrag_mode baseline. This patch (of 4): A subsequent patch will have some order-0 allocations participate in compaction under defrag_mode, to stave off extfrag events. Since this is a sprawling expansion of entry points, and compaction can enter filesystem paths, add lockdep annotations that catches __GFP_FS passing errors. Direct reclaim has had this annotation for a while, and since reclaim and compaction are usually used in conjunction, this is unlikely to unearth old bugs. It's more about future proofing and peace of mind. Link: https://lore.kernel.org/20260722150006.3848560-1-hannes@cmpxchg.org Link: https://lore.kernel.org/20260722150006.3848560-2-hannes@cmpxchg.org Fixes: e3aa7df331bc ("mm: page_alloc: defrag_mode") Signed-off-by: Johannes Weiner <hannes@cmpxchg.org> Reviewed-by: Vlastimil Babka (SUSE) <vbabka@kernel.org> Acked-by: Shakeel Butt <shakeel.butt@linux.dev> Cc: Brendan Jackman <jackmanb@google.com> Cc: Brendan Jackman <brendan.jackman@linux.dev> Cc: David Hildenbrand <david@kernel.org> Cc: Liam R. Howlett <liam@infradead.org> Cc: Lorenzo Stoakes <ljs@kernel.org> Cc: Michal Hocko <mhocko@suse.com> Cc: Mike Rapoport <rppt@kernel.org> Cc: Suren Baghdasaryan <surenb@google.com> Cc: Zi Yan <ziy@nvidia.com> Cc: Brendan Jackman <brendan.jackman@linux.dev> Cc: Gregory Price <gourry@gourry.net> Cc: <stable@vger.kernel.org> Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
2026-08-06hugetlb: evaluate subpool free state while lockedYichong Chen1-2/+4
unlock_or_release_subpool() drops spool->lock before calling subpool_is_free(). However, subpool_is_free() reads fields that are updated under spool->lock, including count, used_hpages and rsv_hpages. Keep the free-state evaluation under the same lock that protects those fields. The reservation accounting and kfree() calls still happen after dropping spool->lock. Link: https://lore.kernel.org/20260721035207.1437935-1-chenyichong@uniontech.com Signed-off-by: Yichong Chen <chenyichong@uniontech.com> Reviewed-by: Joshua Hahn <joshua.hahnjy@gmail.com> Reviewed-by: Jane Chu <jane.chu@oracle.com> Cc: David Hildenbrand <david@kernel.org> Cc: Muchun Song <muchun.song@linux.dev> Cc: Oscar Salvador <osalvador@suse.de> Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
2026-08-06mm/damon: remove trailing semicolons after function definitionsXuewen Wang3-3/+3
Three function definitions terminate with '};' instead of '}', which is unnecessary and inconsistent with kernel coding style: - damon_pa_initcall() in paddr.c - damon_va_initcall() in vaddr.c - damos_get_some_mem_psi_total() in core.c No functional change intended. Link: https://lore.kernel.org/20260721135333.241106-1-sj@kernel.org Signed-off-by: Xuewen Wang <wangxuewen@kylinos.cn> Reviewed-by: SJ Park <sj@kernel.org> Signed-off-by: SJ Park <sj@kernel.org> Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
2026-08-06mm/damon/ops-common: prevent migration fallback to non-target nodesJiahui Zhang1-1/+1
DAMOS_MIGRATE_{HOT,COLD} passes a target NUMA node to migrate_pages(). But alloc_migration_target() only treats mtc->nid as a preferred node unless __GFP_THISNODE is set. Hence target allocation can fall back to another node, and migrate_pages() can report success without placing the folio on the requested target node. Consider a two-node tiered system where node 0 is a fast tier and node 1 is a CPU-less slow tier such as CXL memory, and the user wants to promote hot regions from node 1 to node 0 with a command like: sudo damo start --ops vaddr --target_pid ${workload_pid} \ --damos_action migrate_hot 0 \ --damos_access_rate 70% max Without the __GFP_THISNODE flag, when the memory allocator finds that node 0 is nearly full, it can fall back to node 1 without waking up kswapd. Then the pages allocated for migrate_pages() are still on node 1, and the regions that are expected to be promoted to node 0 are only moved to different physical pages on node 1. Meanwhile, both the mm_migrate_pages tracepoint and DAMOS's own sz_applied statistics (reported via the damos_stat_after_apply_interval tracepoint) show the migrations as successful, which makes the failure practically invisible and hard to investigate. Running a demotion-purpose DAMOS scheme alongside the promotion scheme does not fully avoid this either. If demotion cannot keep up with the promotion rate, allocation can still fall back to node 1 during promotion, and the same misleading statistics show up. Make DAMON's migration target allocation strict by setting __GFP_THISNODE, so that a failed allocation on the target node is reported as a failure instead of silently landing on a different node. This is consistent with alloc_misplaced_dst_folio(), alloc_demote_folio(), and with do_move_pages_to_node(), which all use __GFP_THISNODE for migrations to an explicit destination node. Link: https://lore.kernel.org/20260721135607.251869-1-sj@kernel.org Signed-off-by: Jiahui Zhang <jiahuitry@outlook.com> Reviewed-by: SJ Park <sj@kernel.org> Signed-off-by: SJ Park <sj@kernel.org> Cc: Honggyu Kim <honggyu.kim@sk.com> Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
2026-08-06mm/damon: update outdated comment about DAMOS filter handlingSong Hu1-9/+8
The kernel-doc comment above enum damos_filter_type states that only the anon and memcg type filters are handled by damon_operations (and therefore accounted as 'tried'), and that DAMON_OPS_VADDR and DAMON_OPS_FVADDR do not support those two filter types. Neither is accurate anymore. damos_filter_for_ops() routes every filter type except ADDR and TARGET to the operations layer, and the VADDR and FVADDR operations (the latter being a copy of the former) handle all of those types through damos_folio_filter_match() / damos_va_filter_out(). Update the comment to match the code. Link: https://lore.kernel.org/20260721140011.269802-1-sj@kernel.org Signed-off-by: Song Hu <husong@kylinos.cn> Reviewed-by: SJ Park <sj@kernel.org> Signed-off-by: SJ Park <sj@kernel.org> Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
2026-08-06arm64: remove early_ioremap_reset() call and __late_* macrosSang-Heon Jeon2-5/+0
On arm64, __early_set_fixmap(), __late_set_fixmap() and __late_clear_fixmap() are all __set_fixmap(). Calling early_ioremap_reset() changes nothing. So remove the call and the macros. No functional change. Link: https://lore.kernel.org/20260708170647.362562-4-ekffu200098@gmail.com Signed-off-by: Sang-Heon Jeon <ekffu200098@gmail.com> Acked-by: Will Deacon <will@kernel.org> Cc: Albert Ou <aou@eecs.berkeley.edu> Cc: Alexandre Ghiti <alex@ghiti.fr> Cc: Catalin Marinas <catalin.marinas@arm.com> Cc: David Hildenbrand <david@kernel.org> Cc: Liam R. Howlett <liam@infradead.org> Cc: Lorenzo Stoakes <ljs@kernel.org> Cc: Michal Hocko <mhocko@suse.com> Cc: Mike Rapoport <rppt@kernel.org> Cc: Palmer Dabbelt <palmer@dabbelt.com> Cc: Suren Baghdasaryan <surenb@google.com> Cc: Vlastimil Babka <vbabka@kernel.org> Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
2026-08-06riscv: remove unused __late_set_fixmap() and __late_clear_fixmap()Sang-Heon Jeon1-3/+0
__late_set_fixmap() and __late_clear_fixmap() are only used after early_ioremap_reset() has been called. riscv never calls it because __set_fixmap() works before and after paging_init(). So remove them. No functional change. Link: https://lore.kernel.org/20260708170647.362562-3-ekffu200098@gmail.com Signed-off-by: Sang-Heon Jeon <ekffu200098@gmail.com> Cc: Albert Ou <aou@eecs.berkeley.edu> Cc: Alexandre Ghiti <alex@ghiti.fr> Cc: Catalin Marinas <catalin.marinas@arm.com> Cc: David Hildenbrand <david@kernel.org> Cc: Liam R. Howlett <liam@infradead.org> Cc: Lorenzo Stoakes <ljs@kernel.org> Cc: Michal Hocko <mhocko@suse.com> Cc: Mike Rapoport <rppt@kernel.org> Cc: Palmer Dabbelt <palmer@dabbelt.com> Cc: Suren Baghdasaryan <surenb@google.com> Cc: Vlastimil Babka <vbabka@kernel.org> Cc: Will Deacon <will@kernel.org> Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
2026-08-06mm/early_ioremap: clarify early_ioremap_reset() semanticsSang-Heon Jeon1-3/+7
Patch series "mm/early_ioremap: clarify and clean up early_ioremap_reset()". __late_set_fixmap() and __late_clear_fixmap() are only used after early_ioremap_reset() has been called, but the comment above them does not say anything about that. So arm64, riscv and powerpc, whose __set_fixmap() works before and after paging_init(), describe the same situation in three different ways: calls reset defines the macros arm64 yes yes riscv no yes powerpc no no Patch 1 documents when early_ioremap_reset() needs to be called and that only architectures calling it need to define the macros. Patches 2 and 3 remove the unneeded riscv macros, which are unreachable, and the arm64 reset call and macros, which change nothing. No functional change. This patch (of 3): __late_set_fixmap() and __late_clear_fixmap() are only used after early_ioremap_reset() has been called. arm64, riscv and powerpc all have a __set_fixmap() that works before and after paging_init(), so they do not need to call early_ioremap_reset() or define the macros, but they describe the same situation in three different ways: calls reset defines the macros arm64 yes yes riscv no yes powerpc no no The existing comment is vague and allows all three. Replace it with comments that make it clear when the reset and the macros are needed. No functional change. Link: https://lore.kernel.org/20260708170647.362562-1-ekffu200098@gmail.com Link: https://lore.kernel.org/20260708170647.362562-2-ekffu200098@gmail.com Signed-off-by: Sang-Heon Jeon <ekffu200098@gmail.com> Cc: Albert Ou <aou@eecs.berkeley.edu> Cc: Alexandre Ghiti <alex@ghiti.fr> Cc: Catalin Marinas <catalin.marinas@arm.com> Cc: David Hildenbrand <david@kernel.org> Cc: Liam R. Howlett <liam@infradead.org> Cc: Lorenzo Stoakes <ljs@kernel.org> Cc: Michal Hocko <mhocko@suse.com> Cc: Mike Rapoport <rppt@kernel.org> Cc: Palmer Dabbelt <palmer@dabbelt.com> Cc: Suren Baghdasaryan <surenb@google.com> Cc: Vlastimil Babka <vbabka@kernel.org> Cc: Will Deacon <will@kernel.org> Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
2026-08-06selftests/mm/pagemap_ioctl: fix missing NULL checks after calloc()longlong yan1-1/+7
The pagemap_ioctl selftest allocates memory via calloc() in several places but does not check the return values. If calloc() fails, the subsequent code will dereference a NULL pointer and crash. Additionally, in sanity_tests(), the calloc() failure check incorrectly uses MAP_FAILED (the mmap() error constant) instead of NULL. Since calloc() returns NULL on failure, the check never triggers and a failed allocation goes undetected. Add NULL checks after each calloc() call, and fix the wrong error constant in sanity_tests(). Use ksft_exit_fail_msg() consistent with the existing error handling pattern in the file. Link: https://lore.kernel.org/20260721063611.342-1-yanlonglong@kylinos.cn Signed-off-by: longlong yan <yanlonglong@kylinos.cn> Acked-by: David Hildenbrand (Arm) <david@kernel.org> Reviewed-by: SJ Park <sj@kernel.org> Cc: Shuah Khan <shuah@kernel.org> Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
2026-08-06selftests/mm: use MAP_FAILED for mmap error checklonglong yan1-53/+53
Replace the direct comparison with (void *)-1 with the standard MAP_FAILED macro when checking mmap() Link: https://lore.kernel.org/20260720063439.522-1-yanlonglong@kylinos.cn Signed-off-by: longlong yan <yanlonglong@kylinos.cn> Reviewed-by: Lorenzo Stoakes (ARM) <ljs@kernel.org> Acked-by: David Hildenbrand (Arm) <david@kernel.org> Cc: Liam R. Howlett <liam@infradead.org> Cc: Michal Hocko <mhocko@suse.com> Cc: Mike Rapoport <rppt@kernel.org> Cc: Shuah Khan <shuah@kernel.org> Cc: Suren Baghdasaryan <surenb@google.com> Cc: Vlastimil Babka <vbabka@kernel.org> Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
2026-08-06mm/rmap: batch unmap file folios belonging to uffd-wp VMAsDev Jain1-4/+2
Commit a67fe41e214f ("mm: rmap: support batched unmapping for file large folios") extended batched unmapping for file folios. That also required making pte_install_uffd_wp_if_needed() support batching, but that was left out for the time being. Correctness was maintained by stopping batching if the VMA the folio belongs to is marked uffd-wp. Now that cond_install_uffd_wp_ptes() supports batching, call it with the full batch length and allow folio_unmap_pte_batch() to batch file folios belonging to uffd-wp VMAs. For file folios, if the uffd-wp bit is set, unmapping converts present PTEs into uffd-wp markers. We must ensure that the same PTE range is not reprocessed by the try_to_unmap_one() loop. The page_vma_mapped_walk API ensures this: check_pte() only returns true if any PFN in [pvmw->pfn, pvmw->pfn + nr_pages) is mapped by the PTE. There is no PFN underlying a uffd-wp marker PTE, so check_pte() returns false and the walk skips ahead until it reaches a present entry again. Link: https://lore.kernel.org/20260720065508.2695106-4-dev.jain@arm.com Signed-off-by: Dev Jain <dev.jain@arm.com> Acked-by: David Hildenbrand (Arm) <david@kernel.org> Cc: Anshuman Khandual <anshuman.khandual@arm.com> Cc: Axel Rasmussen <axelrasmussen@google.com> Cc: Barry Song <baohua@kernel.org> Cc: Harry Yoo <harry@kernel.org> Cc: Jann Horn <jannh@google.com> Cc: Kairui Song <kasong@tencent.com> Cc: Lance Yang <lance.yang@linux.dev> Cc: Liam R. Howlett <liam@infradead.org> Cc: Lorenzo Stoakes <ljs@kernel.org> Cc: Michal Hocko <mhocko@suse.com> Cc: Mike Rapoport <rppt@kernel.org> Cc: Rik van Riel <riel@surriel.com> Cc: Ryan Roberts <ryan.roberts@arm.com> Cc: Shakeel Butt <shakeel.butt@linux.dev> Cc: Suren Baghdasaryan <surenb@google.com> Cc: Vlastimil Babka <vbabka@kernel.org> Cc: Wei Xu <weixugc@google.com> Cc: Yuanchu Xie <yuanchu@google.com> Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
2026-08-06mm/memory: batch set uffd-wp markers during zappingDev Jain3-39/+29
Enable batch setting of uffd-wp PTE markers. The code paths passing nr > 1 to zap_install_uffd_wp_if_needed() produce that nr through either folio_pte_batch() or swap_pte_batch(), therefore batching is correct: 1) All PTEs belong to the same type of VMA: anonymous or non-anonymous, wp-armed or non-wp-armed. 2) All PTEs are either marked with uffd-wp or not marked with uffd-wp; the same applies to the pte_swp_uffd_any() check. 3) uffd_supports_wp_marker() is independent of the function parameters. Use set_pte_at() in a loop instead of set_ptes(), because set_ptes() cannot handle nonpresent to nonpresent conversion for nr_pages > 1. Rename the helper to cond_install_uffd_wp_ptes(). Link: https://lore.kernel.org/20260720065508.2695106-3-dev.jain@arm.com Signed-off-by: Dev Jain <dev.jain@arm.com> Acked-by: David Hildenbrand (Arm) <david@kernel.org> Cc: Anshuman Khandual <anshuman.khandual@arm.com> Cc: Axel Rasmussen <axelrasmussen@google.com> Cc: Barry Song <baohua@kernel.org> Cc: Harry Yoo <harry@kernel.org> Cc: Jann Horn <jannh@google.com> Cc: Kairui Song <kasong@tencent.com> Cc: Lance Yang <lance.yang@linux.dev> Cc: Liam R. Howlett <liam@infradead.org> Cc: Lorenzo Stoakes <ljs@kernel.org> Cc: Michal Hocko <mhocko@suse.com> Cc: Mike Rapoport <rppt@kernel.org> Cc: Rik van Riel <riel@surriel.com> Cc: Ryan Roberts <ryan.roberts@arm.com> Cc: Shakeel Butt <shakeel.butt@linux.dev> Cc: Suren Baghdasaryan <surenb@google.com> Cc: Vlastimil Babka <vbabka@kernel.org> Cc: Wei Xu <weixugc@google.com> Cc: Yuanchu Xie <yuanchu@google.com> Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
2026-08-06mm/memory: move pte_install_uffd_wp_if_needed() into memory.cDev Jain3-53/+63
Patch series "Batch unmap of uffd-wp file folios", v2. Currently, batched unmapping is supported if: 1) folio is a file folio, not belonging to uffd-wp VMA 2) folio is anonymous and not swapbacked (lazyfree), not belonging to uffd-wp VMA So the cases which are not supported are 1) folio belonging to uffd-wp VMA 2) folio is anonymous and swapbacked It is easy to see that this adds a lot of cognitive load while reading try_to_unmap_one - we need to remember throughout whether nr_pages == 1 or > 1. The uffd-wp handling in try_to_unmap_one is regarding preserving the uffd-wp state for file folios via pte_install_uffd_wp_if_needed (for anon folio, we handle that while constructing the swap pte). Stop special casing on uffd-wp VMAs by simply adding batching support to pte_install_uffd_wp_if_needed. This patch (of 3): pte_install_uffd_wp_if_needed() has grown too large for mm_inline.h. Move it to memory.c. This helper is only used inside mm/, so declare it in mm/internal.h instead of a public header. While at it, convert the comment to kerneldoc and rename the local arguments from pte/pteval to ptep/pte so the pointer and PTE value are easier to distinguish. Link: https://lore.kernel.org/20260720065508.2695106-1-dev.jain@arm.com Link: https://lore.kernel.org/20260720065508.2695106-2-dev.jain@arm.com Signed-off-by: Dev Jain <dev.jain@arm.com> Acked-by: David Hildenbrand (Arm) <david@kernel.org> Cc: Anshuman Khandual <anshuman.khandual@arm.com> Cc: Axel Rasmussen <axelrasmussen@google.com> Cc: Barry Song <baohua@kernel.org> Cc: Harry Yoo <harry@kernel.org> Cc: Jann Horn <jannh@google.com> Cc: Kairui Song <kasong@tencent.com> Cc: Lance Yang <lance.yang@linux.dev> Cc: Liam R. Howlett <liam@infradead.org> Cc: Lorenzo Stoakes <ljs@kernel.org> Cc: Michal Hocko <mhocko@suse.com> Cc: Mike Rapoport <rppt@kernel.org> Cc: Rik van Riel <riel@surriel.com> Cc: Ryan Roberts <ryan.roberts@arm.com> Cc: Shakeel Butt <shakeel.butt@linux.dev> Cc: Suren Baghdasaryan <surenb@google.com> Cc: Vlastimil Babka <vbabka@kernel.org> Cc: Wei Xu <weixugc@google.com> Cc: Yuanchu Xie <yuanchu@google.com> Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
2026-08-06hugetlb: make hugepage_put_subpool() tolerate NULLYichong Chen2-4/+5
Both callers of hugepage_put_subpool() check whether the subpool pointer is NULL before calling it. Move the NULL check into hugepage_put_subpool() so callers can use the helper unconditionally. This is a follow-up cleanup after using hugepage_put_subpool() from the hugetlbfs_fill_super() failure path. Link: https://lore.kernel.org/20260720073841.1389354-1-chenyichong@uniontech.com Signed-off-by: Yichong Chen <chenyichong@uniontech.com> Reviewed-by: Muchun Song <muchun.song@linux.dev> Cc: David Hildenbrand <david@kernel.org> Cc: Oscar Salvador <osalvador@suse.de> Signed-off-by: Andrew Morton <akpm@linux-foundation.org>