| Age | Commit message (Collapse) | Author | Files | Lines |
|
The checks on num_vfs in pseries_pci_sriov_enable() are ANDed where OR
was apparently intended. Change it to OR.
Fixes: 9a7f6b438664 ("powerpc/pseries/pci: Associate PEs to VFs in configure SR-IOV")
Acked-by: Nayna Jain <nayna@linux.ibm.com>
Tested-by: R Nageswara Sastry <rnsastry@linux.ibm.com>
Cc: stable@vger.kernel.org # 4.16
Signed-off-by: George Wilson <gcwilson@linux.ibm.com>
Signed-off-by: Madhavan Srinivasan <maddy@linux.ibm.com>
|
|
In papr_phy_attest_create_handle(), the params->cmd.length is not
validated before use, which can result in a buffer overlow. Check it and
return -EINVAL if it is either 0 or exceeds sizeof(params->cmd).
Also, params is freed on the success path but not error. Free it on
errors after memory allocation. And free it on negative fd.
Fixes: 86900ab620a4 ("powerpc/pseries: Add a char driver for physical-attestation RTAS")
Acked-by: Haren Myneni <haren@linux.ibm.com>
Acked-by: Nayna Jain <nayna@linux.ibm.com>
Tested-by: R Nageswara Sastry <rnsastry@linux.ibm.com>
Cc: stable@vger.kernel.org # 6.16
Signed-off-by: George Wilson <gcwilson@linux.ibm.com>
Signed-off-by: Madhavan Srinivasan <maddy@linux.ibm.com>
|
|
https://gitlab.freedesktop.org/drm/misc/kernel into drm-next
drm-misc-next for v7.3:
UAPI Changes:
- Remove the default udmabuf size limit of 64MB.
Cross-subsystem Changes:
- Add dmemcg support for eviction, and hook it up for amdgpu and xe.
Core Changes:
- Changes to TTM to be more aggressive when allocating below protection limit!
- Improve dt binding documentation for renesas.
- Add helper to convert physical address back to buddy block,
add that to and improve its kunit test.
Driver Changes:
- Assorted small fixes to ti-sn65dsi86, panthor, imagination, omapdrm,
bridge/synopsys, panel-edp, ssd130x, panel/tdo-tl070wsh30.
- Add Sharp LQ120P1JX51 panel.
- Add dmemcg support to nouveau.
- Various updates and improvements to sun4i, among which YUV and 4k support.
Signed-off-by: Dave Airlie <airlied@redhat.com>
From: Maarten Lankhorst <maarten.lankhorst@linux.intel.com>
Link: https://patch.msgid.link/917d462a-8976-4a15-bec4-4513ec51c5c0@linux.intel.com
|
|
The {begin,end}_current_label_crit_section() has the same issue as the
{__begin,__end} version. That is the check to see if the label has
been updated in the end check forces an unnecessary memory barrier.
We can optimize this the same way we do with the {__begin,__end}
variant by passing in a local variable that carries the state
information from the begin check into the end check.
No functional change.
Signed-off-by: John Johansen <john.johansen@canonical.com>
|
|
Bobby Eshleman says:
====================
net: devmem: allow rx-buf-size > PAGE_SIZE per binding
Every devmem dmabuf binding hands the page_pool PAGE_SIZE niovs today.
On NICs that consume one descriptor per netmem, this caps a single RX
descriptor at PAGE_SIZE and burns CPU on buffer churn.
In this series, we add a bind-time netlink attribute,
NETDEV_A_DMABUF_RX_BUF_SIZE, that lets userspace request a larger niov
size (power of two >= PAGE_SIZE). Drivers must opt in via
queue_mgmt_ops.QCFG_RX_PAGE_SIZE.
Measurements:
Setup: kperf devmem RX/TX cuda, 4 flows, 64 MB messages, 60s, dctcp,
num-rx-queues=4, dmabuf-rx/tx-size-mb=2048, 10 runs per niov size,
mlx5.
niov RX dev Gbps RX flow avg Gbps app sys %
----- ---------------- ----------------- ----------------
4K 300.63 +/- 53.21 75.16 +/- 13.30 54.15 +/- 10.23
16K 321.35 +/- 28.20 80.34 +/- 7.05 41.05 +/- 8.87
32K 347.63 +/- 2.20 86.91 +/- 0.55 44.54 +/- 3.51
64K 332.11 +/- 14.26 83.03 +/- 3.56 35.47 +/- 3.11
RX app sys % drops ~19% from 4K to 64K.
kperf support (not yet merged):
https://github.com/facebookexperimental/kperf/commit/8837577f920876bce6986ec18869ac04439ebcd2
====================
Link: https://patch.msgid.link/20260805-tcpdm-large-niovs-v8-0-3e0225e2808c@meta.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
Add a new devmem test case for binding the dmabuf with rx-page-size=16K.
The test sweeps RX payload sizes straddling the niov boundary to cover
the sub-niov, exact-niov, and multi-niov RX paths.
Silence pylint invalid-name (`with open() as f`) and too-many-arguments
(ncdevmem_rx grew to 6 args) at file scope.
Acked-by: Stanislav Fomichev <sdf@fomichev.me>
Reviewed-by: Nikolay Aleksandrov <razor@blackwall.org>
Signed-off-by: Bobby Eshleman <bobbyeshleman@meta.com>
Link: https://patch.msgid.link/20260805-tcpdm-large-niovs-v8-3-3e0225e2808c@meta.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
Add -b <bytes> to request a non-default niov size via
NETDEV_A_DMABUF_RX_PAGE_SIZE. When the value exceeds PAGE_SIZE,
udmabuf_alloc() switches to an MFD_HUGETLB-backed memfd so each 2 MB
hugepage produces one naturally-aligned sg entry.
Acked-by: Stanislav Fomichev <sdf@fomichev.me>
Reviewed-by: Nikolay Aleksandrov <razor@blackwall.org>
Signed-off-by: Bobby Eshleman <bobbyeshleman@meta.com>
Link: https://patch.msgid.link/20260805-tcpdm-large-niovs-v8-2-3e0225e2808c@meta.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
Every devmem dmabuf binding today hands the page_pool PAGE_SIZE niovs.
This caps a single RX descriptor at PAGE_SIZE, burning CPU on buffer
churn for large flows.
Add a bind-time netlink attribute, NETDEV_A_DMABUF_RX_PAGE_SIZE, that
lets userspace request a larger niov size. The value must be a power of
two >= PAGE_SIZE.
The TX path is changed to always pass PAGE_SIZE.
Measurements:
Setup: kperf in devmem RX/TX cuda mode, 4 flows, 64 MB messages, 60s,
dctcp, num-rx-queues=4, dmabuf-rx/tx-size-mb=2048, 10 runs per niov
size, mlx5.
CPU Util:
niov net sirq % net idle % app sys % app idle %
----- ---------------- ---------------- ---------------- ----------------
4K 62.38 +/- 8.27 33.40 +/- 7.51 54.15 +/- 10.23 43.67 +/- 10.53
16K 58.91 +/- 5.35 35.23 +/- 5.88 41.05 +/- 8.87 56.42 +/- 9.24
32K 64.12 +/- 0.68 31.09 +/- 1.48 44.54 +/- 3.51 52.63 +/- 3.65
64K 54.69 +/- 5.54 39.67 +/- 5.81 35.47 +/- 3.11 61.97 +/- 3.27
RX app sys % drops ~19% from 4K to 64K.
Throughput:
niov RX dev Gbps RX flow avg Gbps
----- ---------------- -----------------
4K 300.63 +/- 53.21 75.16 +/- 13.30
16K 321.35 +/- 28.20 80.34 +/- 7.05
32K 347.63 +/- 2.20 86.91 +/- 0.55
64K 332.11 +/- 14.26 83.03 +/- 3.56
Throughput seems to increase, but the stdev is pretty wide so could just
be noise.
kperf support (not yet merged):
https://github.com/facebookexperimental/kperf/commit/8837577f920876bce6986ec18869ac04439ebcd2
Acked-by: Stanislav Fomichev <sdf@fomichev.me>
Reviewed-by: Mina Almasry <almasrymina@google.com>
Reviewed-by: Nikolay Aleksandrov <razor@blackwall.org>
Signed-off-by: Bobby Eshleman <bobbyeshleman@meta.com>
Link: https://patch.msgid.link/20260805-tcpdm-large-niovs-v8-1-3e0225e2808c@meta.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
To support up to 8 packets per CQE, update related CQE processing
code and structures.
Update ethtool handlers to set this feature.
Update per queue stat to show the coalesced CQE counters.
This feature is supported on NIC hardware showing the relevant
PF flag.
Signed-off-by: Haiyang Zhang <haiyangz@microsoft.com>
Reviewed-by: Simon Horman <horms@kernel.org>
Reviewed-by: Breno Leitao <leitao@debian.org>
Link: https://patch.msgid.link/20260805185404.1052177-1-haiyangz@linux.microsoft.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
ovs_packet spec needs to fetch the right uAPI header,
like other ovs specs already do. Otherwise build breaks
on very old distros (Ubuntu 22.04).
Spec was added by commit b82bfddc46e2 ("netlink: specs: add OVS packet
family specification").
Reviewed-by: Hangbin Liu <liuhangbin@kylinos.cn>
Link: https://patch.msgid.link/20260807001918.61957-1-kuba@kernel.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
In uapi/asm/hwprobe.h file, (1 << N) is used to define the bit field
which causes checkpatch to warn. Use _BITUL(N) and _BITULL(N) to avoid
these warnings.
Signed-off-by: Jesse Taube <jtaubepe@redhat.com>
Link: https://patch.msgid.link/20260805153346.1036988-1-jtaubepe@redhat.com
[pjw@kernel.org: updated to apply and to cover new extensions]
Signed-off-by: Paul Walmsley <pjw@kernel.org>
|
|
The vDSO for compat processes can also contain alternative entries.
Patch those, too.
Signed-off-by: Thomas Weißschuh <thomas.weissschuh@linutronix.de>
Reviewed-by: Nam Cao <namcao@linutronix.de>
Link: https://patch.msgid.link/20260630-riscv-vdso32-alternative-v1-3-a32fd89b7b1c@linutronix.de
Signed-off-by: Paul Walmsley <pjw@kernel.org>
|
|
Currently the alternative sections are extracted from the vDSO
binaries at runtime. This has runtime overhead and also doesn't
work for the compat vDSO.
Use the offsets generated during the build instead, fixing both issues.
Signed-off-by: Thomas Weißschuh <thomas.weissschuh@linutronix.de>
Reviewed-by: Nam Cao <namcao@linutronix.de>
Link: https://patch.msgid.link/20260630-riscv-vdso32-alternative-v1-2-a32fd89b7b1c@linutronix.de
Signed-off-by: Paul Walmsley <pjw@kernel.org>
|
|
Currently the alternative sections are extracted from the vDSO
binaries at runtime. This has runtime overhead and also doesn't
work for the compat vDSO.
Extract the offsets of the alternative section during the build with the
existing symbol extraction machinery.
Also add dummy symbols as fallback.
Signed-off-by: Thomas Weißschuh <thomas.weissschuh@linutronix.de>
Reviewed-by: Nam Cao <namcao@linutronix.de>
Link: https://patch.msgid.link/20260630-riscv-vdso32-alternative-v1-1-a32fd89b7b1c@linutronix.de
Signed-off-by: Paul Walmsley <pjw@kernel.org>
|
|
Add Ziccamoa, Ziccif, and Za64rs to riscv_isa_ext[] so they can be
parsed from devicetree/ACPI ISA strings. Ziccrse is already present
in cpufeature; this patch only adds its hwprobe exposure.
Expose all four extensions via hwprobe through new bits in
RISCV_HWPROBE_KEY_IMA_EXT_1 (RISCV_HWPROBE_EXT_ZICCAMOA, _ZICCIF,
_ZICCRSE, _ZA64RS), so userspace can probe each of these
RVA23U64-mandatory extensions individually.
Reviewed-by: Jesse Taube <jtaubepe@redhat.com>
Signed-off-by: Andrew Jones <andrew.jones@oss.qualcomm.com>
Signed-off-by: Guodong Xu <docular.xu@gmail.com>
Link: https://patch.msgid.link/20260701-rva23u64-hwprobe-v2-v5-6-2c61f94a695a@gmail.com
[pjw@kernel.org: updated to apply]
Signed-off-by: Paul Walmsley <pjw@kernel.org>
|
|
Zicclsm requires misaligned support for all regular load and store
instructions, both scalar and vector, but not AMOs or other
specialized forms of memory access, to main memory regions with both
the cacheability and coherence PMAs, as defined in the profiles spec.
Even though mandated, misaligned loads and stores might execute
extremely slowly. Standard software distributions should assume their
existence only for correctness, not for performance.
Reviewed-by: Conor Dooley <conor.dooley@microchip.com>
Reviewed-by: Andy Chiu <andy.chiu@sifive.com>
Reviewed-by: Charlie Jenkins <charlie@rivosinc.com>
Tested-by: Charlie Jenkins <charlie@rivosinc.com>
Signed-off-by: Jesse Taube <jesse@rivosinc.com>
[andrew.jones: Rebased, rewrote doc text, minor commit message revisions]
Signed-off-by: Andrew Jones <andrew.jones@oss.qualcomm.com>
Signed-off-by: Guodong Xu <docular.xu@gmail.com>
Link: https://patch.msgid.link/20260701-rva23u64-hwprobe-v2-v5-5-2c61f94a695a@gmail.com
[pjw@kernel.org: updated to apply; added Andrew's username to his tag comments]
Signed-off-by: Paul Walmsley <pjw@kernel.org>
|
|
Specify that chapter 27 refers to version 20191213 of the RISC-V ISA
Unprivileged Architecture. The chapter numbering differs across
specification versions - for example, in version 20250508, the ISA
Extension Naming Conventions is chapter 36, not chapter 27.
Historical versions of the RISC-V specification can be found via Link [1].
Acked-by: Conor Dooley <conor.dooley@microchip.com>
Link: https://riscv.org/specifications/ratified/ [1]
Fixes: 99e2266f2460 ("RISC-V: clarify ISA string ordering rules in cpu.c")
Signed-off-by: Guodong Xu <guodong@riscstar.com>
Link: https://patch.msgid.link/20260125-supm-ext-id-v2-3-1e3b9714c860@riscstar.com
Signed-off-by: Paul Walmsley <pjw@kernel.org>
|
|
The base extensions are often lowercase and were written as lowercase in
hwcap, but other references to these extensions in the kernel are
uppercase. Standardize the case to make it easier to handle macro
expansion.
Acked-by: Anup Patel <anup@brainfault.org>
Reviewed-by: Anup Patel <anup@brainfault.org>
Signed-off-by: Charlie Jenkins <thecharlesjenkins@gmail.com>
[andrew.jones: Apply KVM_ISA_EXT_ARR(), fixup all KVM use.]
Signed-off-by: Andrew Jones <andrew.jones@oss.qualcomm.com>
Signed-off-by: Guodong Xu <docular.xu@gmail.com>
Link: https://patch.msgid.link/20260701-rva23u64-hwprobe-v2-v5-4-2c61f94a695a@gmail.com
[pjw@kernel.org: fixed a checkpatch warning; added Andrew's username to his tag comments]
Signed-off-by: Paul Walmsley <pjw@kernel.org>
|
|
The ftrace selftest multiple_kprobes.tc registers kprobe events on
the first 256 text symbols from /proc/kallsyms. If handle_break() is
selected as a probe target on RISC-V, the breakpoint exception path
can trap again before it reaches the kprobe breakpoint handler. That
recursively enters do_trap_break() and can make the system
unresponsive.
Mark handle_break() and its local probe dispatch helpers as nokprobe
symbols so they are added to the kprobe blacklist, matching other
low-level breakpoint exception paths.
Signed-off-by: Rui Qi <qirui.001@bytedance.com>
Reviewed-by: Nam Cao <namcao@linutronix.de>
Link: https://patch.msgid.link/20260806043807.2583001-1-qirui.001@bytedance.com
Signed-off-by: Paul Walmsley <pjw@kernel.org>
|
|
Instead of adding these to -march globally, use .option arch to use
instructions from these extensions only in code paths where we know they
are available, like how it is done for most other extensions.
TOOLCHAIN_HAS_{ZACAS,ZABHA} already depend on AS_HAS_OPTION_ARCH, so
this is not a functionality regression even on older assemblers.
Although the compiler is unlikely to generate atomics on its own accord,
this aligns handling of Zacas and Zabha with most other extensions and
improves consistency on how assembly code requiring extra extensions is
written in kernel code.
This is analogous to the use of __LSE_PREAMBLE or .arch_extension lse in
arm64 code.
Signed-off-by: Vivian Wang <wangruikang@iscas.ac.cn>
Reviewed-by: Jesse Taube <jtaubepe@redhat.com>
Link: https://patch.msgid.link/20260717-riscv-no-zacas-zabha-in-march-v1-1-82b5b0799fb6@iscas.ac.cn
Signed-off-by: Paul Walmsley <pjw@kernel.org>
|
|
Commit 4785aa802853 ("cpuidle, ACPI: Evaluate LPI arch_flags for
broadcast timer") replaced the generic nonzero check for LPI
architectural context loss flags with arch_get_idle_state_flags().
RISC-V does not implement the helper, so it falls back to the stub
that returns 0. Consequently, CPUIDLE_FLAG_TIMER_STOP is not set when
an LPI state loses the hart timer context, preventing cpuidle from
using a broadcast timer for that state.
Implement the RISC-V helper and map the hart timer context loss flag
to CPUIDLE_FLAG_TIMER_STOP.
Fixes: 4785aa802853 ("cpuidle, ACPI: Evaluate LPI arch_flags for broadcast timer")
Cc: stable@vger.kernel.org
Acked-by: Sudeep Holla <sudeep.holla@kernel.org>
Reviewed-by: Yixun Lan <dlan@kernel.org>
Reviewed-by: Sunil V L <sunilvl@oss.qualcomm.com>
Reviewed-by: Huisong Li <lihuisong@huawei.com>
Signed-off-by: Peixin Xie <peixin.xie@linux.spacemit.com>
Link: https://patch.msgid.link/20260803-riscv-acpi-lpi-timer-v3-1-520fa13732f5@linux.spacemit.com
Signed-off-by: Paul Walmsley <pjw@kernel.org>
|
|
After commit 9b3a2be84803 ("riscv: Remove support for XIP kernel"),
something relatd with XIP are still there. Remove them to clean up the
code.
Signed-off-by: Jisheng Zhang <jszhang@kernel.org>
Reviewed-by: Nam Cao <namcao@linutronix.de>
Link: https://patch.msgid.link/20260805010049.13918-1-jszhang@kernel.org
Signed-off-by: Paul Walmsley <pjw@kernel.org>
|
|
Add support for the srmcfg CSR defined in the Ssqosid ISA extension.
The CSR contains two fields:
- Resource Control ID (RCID) for resource allocation
- Monitoring Counter ID (MCID) for tracking resource usage
Requests from a hart to shared resources are tagged with these IDs,
allowing resource usage to be associated with the running task.
Add a srmcfg field to thread_struct with the same format as the CSR.
The context-switch path writes the field to the CSR, and
resctrl_arch_set_closid_rmid() updates it when a task is assigned to a
resctrl control or monitoring group.
A per-cpu cpu_srmcfg_default holds the default srmcfg for each CPU, set
by resctrl_arch_set_cpu_default_closid_rmid() on CPU group assignment.
On context switch, RCID and MCID inherit from the CPU default
independently: a task whose thread RCID field is zero takes the CPU
default's RCID, and likewise for MCID.
A per-cpu cpu_srmcfg variable mirrors the CSR state to avoid
redundant writes. L1D-hot memory access is faster than a CSR read and
avoids traps under virtualization.
Link: https://github.com/riscv/riscv-ssqosid/releases/tag/v1.0
Assisted-by: Claude:claude-opus-4-7
Co-developed-by: Kornel Dulęba <mindal@semihalf.com>
Signed-off-by: Kornel Dulęba <mindal@semihalf.com>
Signed-off-by: Drew Fustini <fustini@kernel.org>
Link: https://patch.msgid.link/20260729-dfustini-atl-sc-cbqri-dt-v6-3-7c22b05d461b@kernel.org
Signed-off-by: Paul Walmsley <pjw@kernel.org>
|
|
Ssqosid is the RISC-V Quality-of-Service (QoS) Identifiers specification
which defines the Supervisor Resource Management Configuration (srmcfg)
register.
Link: https://github.com/riscv/riscv-ssqosid/releases/tag/v1.0
Co-developed-by: Kornel Dulęba <mindal@semihalf.com>
Signed-off-by: Kornel Dulęba <mindal@semihalf.com>
Signed-off-by: Drew Fustini <fustini@kernel.org>
Link: https://patch.msgid.link/20260729-dfustini-atl-sc-cbqri-dt-v6-2-7c22b05d461b@kernel.org
[pjw@kernel.org: updated to apply]
Signed-off-by: Paul Walmsley <pjw@kernel.org>
|
|
Document the ratified Supervisor-mode Quality of Service ID (Ssqosid)
extension v1.0.
Link: https://github.com/riscv/riscv-ssqosid/releases/tag/v1.0
Acked-by: Conor Dooley <conor.dooley@microchip.com>
Signed-off-by: Drew Fustini <fustini@kernel.org>
Link: https://patch.msgid.link/20260729-dfustini-atl-sc-cbqri-dt-v6-1-7c22b05d461b@kernel.org
Signed-off-by: Paul Walmsley <pjw@kernel.org>
|
|
Add description for the Smcdeleg/Ssccfg extension.
Signed-off-by: Atish Patra <atishp@rivosinc.com>
Acked-by: Conor Dooley <conor.dooley@microchip.com>
Link: https://patch.msgid.link/20260807-counter_delegation-v9-10-58658104e487@meta.com
Signed-off-by: Paul Walmsley <pjw@kernel.org>
|
|
Smcdeleg extension allows the M-mode to delegate selected counters
to S-mode so that it can access those counters and correpsonding
hpmevent CSRs without M-mode.
Ssccfg (‘Ss’ for Privileged architecture and Supervisor-level
extension, ‘ccfg’ for Counter Configuration) provides access to
delegated counters and new supervisor-level state.
This patch just enables these definitions and enable parsing.
Signed-off-by: Atish Patra <atishp@rivosinc.com>
Reviewed-by: Charlie Jenkins <thecharlesjenkins@gmail.com>
Link: https://patch.msgid.link/20260807-counter_delegation-v9-9-58658104e487@meta.com
Signed-off-by: Paul Walmsley <pjw@kernel.org>
|
|
This adds the scountinhibit CSR definition and S-mode accessible hpmevent
bits defined by smcdeleg/ssccfg. scountinhibit allows S-mode to start/stop
counters directly from S-mode without invoking SBI calls to M-mode. It is
also used to figure out the counters delegated to S-mode by the M-mode as
well.
Signed-off-by: Kaiwen Xue <kaiwenx@rivosinc.com>
Reviewed-by: Clément Léger <cleger@rivosinc.com>
Reviewed-by: Charlie Jenkins <thecharlesjenkins@gmail.com>
Tested-by: Charlie Jenkins <thecharlesjenkins@gmail.com>
[pjw@kernel.org: fixed subject typo]
Link: https://patch.msgid.link/20260807-counter_delegation-v9-8-58658104e487@meta.com
Signed-off-by: Paul Walmsley <pjw@kernel.org>
|
|
Add the description for Smcntrpmf ISA extension
Acked-by: Rob Herring (Arm) <robh@kernel.org>
Signed-off-by: Atish Patra <atishp@rivosinc.com>
Link: https://patch.msgid.link/20260807-counter_delegation-v9-7-58658104e487@meta.com
Signed-off-by: Paul Walmsley <pjw@kernel.org>
|
|
Smcntrpmf extension allows M-mode to enable privilege mode filtering
for cycle/instret counters. However, the cyclecfg/instretcfg CSRs are
available in Ssccfg only if Smcntrpmf is present.
That's why, kernel needs to detect presence of Smcntrpmf extension and
enable privilege mode filtering for cycle/instret counters.
Reviewed-by: Clément Léger <cleger@rivosinc.com>
Signed-off-by: Atish Patra <atishp@rivosinc.com>
Reviewed-by: Charlie Jenkins <thecharlesjenkins@gmail.com>
Tested-by: Charlie Jenkins <thecharlesjenkins@gmail.com>
Link: https://patch.msgid.link/20260807-counter_delegation-v9-6-58658104e487@meta.com
Signed-off-by: Paul Walmsley <pjw@kernel.org>
|
|
The indirect CSR requires multiple instructions to read/write CSR.
Add a few helper macros for ease of usage.
These have to be macros rather than functions. csr_read()/csr_write()
stringify their CSR argument into the inline asm template via
__ASM_STR(), so the CSR number must be a literal token; passing it as a
function parameter emits "csrr %0, iregcsr", which the assembler rejects
with "unknown CSR `iregcsr'". The stringification happens in the
preprocessor, before inlining or constant propagation, so it cannot be
worked around by forcing inlining or by only ever passing constants -
gcc 12, gcc 16 and clang 22 all reject it alike. Underneath, csrr/csrw
encode the CSR as a 12-bit immediate and RISC-V has no register-indirect
form, which is also why asm/csr.h keeps every one of its accessors as a
macro.
Signed-off-by: Atish Patra <atishp@rivosinc.com>
Reviewed-by: Charlie Jenkins <thecharlesjenkins@gmail.com>
Tested-by: Charlie Jenkins <thecharlesjenkins@gmail.com>
[pjw@kernel.org: expand "ind" abbreviation]
Link: https://patch.msgid.link/20260807-counter_delegation-v9-5-58658104e487@meta.com
Signed-off-by: Paul Walmsley <pjw@kernel.org>
|
|
Add the S[m|s]csrind ISA extension description.
Acked-by: Rob Herring (Arm) <robh@kernel.org>
Signed-off-by: Atish Patra <atishp@rivosinc.com>
[pjw@kernel.org: use official extension names in the patch description]
Link: https://patch.msgid.link/20260807-counter_delegation-v9-4-58658104e487@meta.com
Signed-off-by: Paul Walmsley <pjw@kernel.org>
|
|
The S[m|s]csrind extensions extend the indirect CSR access mechanism
defined in Smaia/Ssaia extensions.
This patch just enables the definition and parsing.
Signed-off-by: Atish Patra <atishp@rivosinc.com>
Reviewed-by: Charlie Jenkins <thecharlesjenkins@gmail.com>
Tested-by: Charlie Jenkins <thecharlesjenkins@gmail.com>
[pjw@kernel.org: use official RISC-V extension names in the patch description]
Link: https://patch.msgid.link/20260807-counter_delegation-v9-3-58658104e487@meta.com
Signed-off-by: Paul Walmsley <pjw@kernel.org>
|
|
This adds definitions of new CSRs and bits defined in the Smcsrind and
Sscsrind ISA extensions. These CSRs enable the indirect CSR accesses
mechanism to access any indirect CSRs in M-, S-, and VS-mode. The
range of the select values and ireg will be defined by the ISA
extension that are based on the Smcsrind and Sscsrind extensions.
Signed-off-by: Kaiwen Xue <kaiwenx@rivosinc.com>
Reviewed-by: Clément Léger <cleger@rivosinc.com>
Signed-off-by: Atish Patra <atishp@rivosinc.com>
Reviewed-by: Charlie Jenkins <thecharlesjenkins@gmail.com>
Tested-by: Charlie Jenkins <thecharlesjenkins@gmail.com>
[pjw@kernel.org: clean up the patch description; use official RISC-V extension names]
Link: https://patch.msgid.link/20260807-counter_delegation-v9-2-58658104e487@meta.com
Signed-off-by: Paul Walmsley <pjw@kernel.org>
|
|
Sashiko pointed out various UAF and memory leak issues around
pmu_sbi_device_probe() error paths.
If the probe fails, here are list of cleanups needed.
a. Already registered pmu must be freed
b. per cpu IRQ must be released
c. pmu_ctr_list data structure must be freed
d. cpu hotplug state must be cleaned up only if added.
Fix the resource cleanup by reorganizing the code around probe failure.
Reported-by: Sashiko AI <sashiko-bot@kernel.org>
Reviewed-by: Charlie Jenkins <thecharlesjenkins@gmail.com>
Signed-off-by: Atish Patra <atishp@meta.com>
Link: https://patch.msgid.link/20260807-counter_delegation-v9-1-58658104e487@meta.com
Signed-off-by: Paul Walmsley <pjw@kernel.org>
|
|
This reverts commit 5d15d2ad36b0 ("riscv: hwprobe: Fix stale vDSO data for
late-initialized keys at boot"). The commit ensures synchronization between
the unaligned vector access speed probe kthread and vDSO data read. But now
that the kthread has been removed, this commit can be reverted.
Signed-off-by: Nam Cao <namcao@linutronix.de>
Tested-by: Anirudh Srinivasan <asrinivasan@oss.tenstorrent.com>
Link: https://patch.msgid.link/50ca78a649faf53f8941bc94c9cf8268b3644d38.1781666867.git.namcao@linutronix.de
Signed-off-by: Paul Walmsley <pjw@kernel.org>
|
|
A kthread is used to run check_vector_unaligned_access() to optimize boot
time, allowing the kernel to continue booting without waiting for the
unaligned vector speed probe to finish.
However, this asynchronous approach introduces several complications.
First, the kthread may not complete before a user reads vDSO data,
resulting in incorrect values. This was previously addressed by
commit 5d15d2ad36b0 ("riscv: hwprobe: Fix stale vDSO data for
late-initialized keys at boot"), which added complex synchronization
between the kthread and vDSO reads.
Second, it was discovered that the kthread may not finish before
vec_check_unaligned_access_speed_all_cpus() (marked with __init) is freed,
triggering a page fault.
These issues raise the question of whether the kthread is worth the added
complexity. A past boot time regression report was actually unrelated to
synchronous probing; it was caused by the probe running serially. Since
switching to a parallel probe, no further complaints have been made.
Furthermore, the unaligned scalar access speed probe takes the same amount
of time, runs synchronously, and has caused no issues.
Testing shows no noticeable boot time slowdown when running the vector
probe synchronously (0.464474s with kthread vs. 0.457991s without).
Remove the kthread usage and run the probe synchronously. This simplifies
the boot flow and allows for the revert of commit 5d15d2ad36b0 ("riscv:
hwprobe: Fix stale vDSO data for late-initialized keys at boot")
Reported-by: Anirudh Srinivasan <asrinivasan@oss.tenstorrent.com>
Closes: https://lore.kernel.org/linux-riscv/20260612-vec_unaligned_drop_init-v1-1-df969210ae34@oss.tenstorrent.com/
Fixes: e7c9d66e313b ("RISC-V: Report vector unaligned access speed hwprobe")
Cc: stable@vger.kernel.org
Signed-off-by: Nam Cao <namcao@linutronix.de>
Acked-by: Jesse Taube <jtaubepe@redhat.com>
Tested-by: Anirudh Srinivasan <asrinivasan@oss.tenstorrent.com>
Link: https://patch.msgid.link/1c378963f27c5960e8a57c50b8b444d30954cb54.1781666867.git.namcao@linutronix.de
[pjw@kernel.org: updated to apply; adjusted Fixes: tag; fixed my own manual patch application error]
Signed-off-by: Paul Walmsley <pjw@kernel.org>
|
|
Since default_power_off() never returns, annotate it with the __noreturn
attribute to improve compiler optimizations.
Signed-off-by: Thorsten Blum <thorsten.blum@linux.dev>
Link: https://patch.msgid.link/20260727100339.410466-2-thorsten.blum@linux.dev
Signed-off-by: Paul Walmsley <pjw@kernel.org>
|
|
Move the dma_contiguous_reserve() call from setup_bootmem() to
misc_mem_init(), placing it after arch_numa_init(). This ensures that
NUMA topology is initialized when reserving contiguous memory for DMA.
Tested with CMA_SIZE_PERNUMA enabled and two NUMA nodes on QEMU:
qemu-system-riscv64 \
-machine virt \
-nographic \
-smp 2 -m 512M \
-numa node,nodeid=0,cpus=0,memdev=m0 \
-numa node,nodeid=1,cpus=1,memdev=m1 \
-object memory-backend-ram,id=m0,size=256M \
-object memory-backend-ram,id=m1,size=256M \
-kernel arch/riscv/boot/Image \
-append "console=ttyS0 earlycon loglevel=7 cma=16M"
Unpatched kernel's log (one global CMA pool and no per-nodepools, 16 MiB
cma-reserved):
[ 0.000000] cma: Reserved 16 MiB at 0x000000009ee00000
...
[ 0.000000] Initmem setup node 0 [mem 0x0000000080000000-0x000000008fffffff]
[ 0.000000] Initmem setup node 1 [mem 0x0000000090000000-0x000000009fffffff]
...
[ 0.056397] smp: Brought up 2 nodes, 2 CPUs
[ 0.066135] Memory: 449288K/524288K available (12279K kernel code, 5980K rwdata, 6144K rodata, 2467K init, 482K bss, 54504K reserved, 16384K cma-reserved)
Patched kernel's log (three CMA pools are created, 48 MiB total),
[ 0.000000] cma: Reserved 16 MiB at 0x000000009ee00000
[ 0.000000] cma: Reserved 16 MiB at 0x000000008ee00000
[ 0.000000] cma: Reserved 16 MiB at 0x000000009de00000
...
[ 0.054251] smp: Brought up 2 nodes, 2 CPUs
[ 0.064082] Memory: 416520K/524288K available (12279K kernel code, 5980K rwdata, 6144K rodata, 2467K init, 482K bss, 54504K reserved, 49152K cma-reserved)
Signed-off-by: Eder Zulian <ezulian@redhat.com>
Link: https://patch.msgid.link/20260714185648.1082483-1-ezulian@redhat.com
Signed-off-by: Paul Walmsley <pjw@kernel.org>
|
|
Change the shadow stack size calculation from RLIMIT_STACK/2 (capped at
2GB) to RLIMIT_STACK/8 (capped at 512MB), following David Laight's
analysis and recommendation.
Rationale:
David Laight pointed out that the focus should be on the ratio between
shadow stack size and the normal stack size, rather than just the
absolute upper limit. His analysis showed that while there are many
functions with small stack frames, the majority have stack deltas of
over 64 bytes due to saved registers and local variables.
Shadow stacks only store return addresses (8 bytes per entry on 64-bit
systems), whereas normal stack frames typically consume 64+ bytes. This
8:64 byte ratio means that programs using a lot of stack space are
dominated by large buffer allocations and local variables, not extreme
recursion depths with minimal local data.
For example, with the default RLIMIT_STACK of 8MB:
- RLIMIT_STACK/2 gives a 4MB shadow stack supporting 512K nested calls
- RLIMIT_STACK/8 gives a 1MB shadow stack supporting 128K nested calls
Given typical stack frame sizes of 64+ bytes, RLIMIT_STACK/8 is still
conservative and provides adequate depth for practical applications.
David noted that this could even be safely halved again.
This reduction also better accommodates memory-constrained platforms.
On systems with limited physical memory, allocating large shadow stacks
can cause virtual memory allocation failures when overcommit mode is set
to OVERCOMMIT_GUESS or OVERCOMMIT_NEVER.
Suggested-by: David Laight <david.laight.linux@gmail.com>
Link: https://lore.kernel.org/all/20260518105725.7afe7a4c@pumpkin/
Signed-off-by: Zong Li <zong.li@sifive.com>
Link: https://patch.msgid.link/20260522093634.3530233-1-zong.li@sifive.com
Signed-off-by: Paul Walmsley <pjw@kernel.org>
|
|
Listing all flags for each object file is tedious and error-prone.
Replace it with a simpler solution.
Link: https://lore.kernel.org/all/20260630135316-f26f0e0f-c08c-4d4d-9963-10f9985a7689@linutronix.de/
Signed-off-by: Thomas Weißschuh <thomas.weissschuh@linutronix.de>
Link: https://patch.msgid.link/20260701-riscv-vdso-lto-v1-2-89db0cd82077@linutronix.de
Signed-off-by: Paul Walmsley <pjw@kernel.org>
|
|
Use Svinval in update_mmu_cache_range() when the extension is available.
Signed-off-by: Xu Lu <luxu.kernel@bytedance.com>
Link: https://patch.msgid.link/20260715132009.10634-3-luxu.kernel@bytedance.com
Tested-by: Klara Modin <klarasmodin@gmail.com>
Signed-off-by: Paul Walmsley <pjw@kernel.org>
|
|
Only flush TLB entries for the specified mm in update_mmu_cache_range().
Signed-off-by: Xu Lu <luxu.kernel@bytedance.com>
Link: https://patch.msgid.link/20260715132009.10634-2-luxu.kernel@bytedance.com
Signed-off-by: Paul Walmsley <pjw@kernel.org>
|
|
Add a test case validating that kprobes correctly simulates the
c.jal instruction on RV32.
The test uses two probe points: a forward c.jal and a backward
c.jal, and verifies that the containing function returns the
expected magic value KPROBE_TEST_MAGIC after kprobe interception.
Co-developed-by: Xiaofeng Yuan <xiaofengmian@163.com>
Signed-off-by: Nam Cao <namcao@linutronix.de>
Signed-off-by: Xiaofeng Yuan <xiaofengmian@163.com>
Reviewed-by: Charlie Jenkins <thecharlesjenkins@gmail.com>
Tested-by: Charlie Jenkins <thecharlesjenkins@gmail.com>
Link: https://patch.msgid.link/20260701081033.49871-3-xiaofengmian@163.com
Signed-off-by: Paul Walmsley <pjw@kernel.org>
|
|
The c.jal instruction is currently marked REJECTED in kprobes
instruction decoding, but it should be SIMULATED like other
compressed jump instructions.
Add simulate_c_jal() which saves the return address to RA and
sets the program counter to the target offset, reusing
simulate_c_j for the common jump logic.
Although c.jal is RV32-only, the function compiles unconditionally.
On RV64, riscv_insn_is_c_jal() always returns 0, so the simulation
code is never invoked and the small overhead in kernel size is
acceptable.
Signed-off-by: Xiaofeng Yuan <xiaofengmian@163.com>
Reviewed-by: Charlie Jenkins <thecharlesjenkins@gmail.com>
Tested-by: Charlie Jenkins <thecharlesjenkins@gmail.com>
Reviewed-by: Nam Cao <namcao@linutronix.de>
Tested-by: Nam Cao <namcao@linutronix.de>
Link: https://patch.msgid.link/20260701081033.49871-2-xiaofengmian@163.com
Signed-off-by: Paul Walmsley <pjw@kernel.org>
|
|
Implement the various required hooks and enable
ARCH_HAS_ACPI_TABLE_UPGRADE to allow use for ACPI_TABLE_UPGRADE, which
is useful for debugging ACPI table problems.
The implementation is based on arm64's of the same feature due to the
similarities of the requirements of the two platforms.
Signed-off-by: Vivian Wang <wangruikang@iscas.ac.cn>
Link: https://patch.msgid.link/20260616-riscv-acpi-table-upgrade-v1-2-45902d2dedf9@iscas.ac.cn
[pjw@kernel.org: updated to apply]
Signed-off-by: Paul Walmsley <pjw@kernel.org>
|
|
drivers/acpi/tables.c uses NR_FIX_BTMAPS without including
<asm/fixmap.h>. This isn't a problem for existing archs, but would be
when ARCH_HAS_ACPI_TABLE_UPGRADE is enabled for RISC-V. Add the missing
include.
Signed-off-by: Vivian Wang <wangruikang@iscas.ac.cn>
Link: https://patch.msgid.link/20260616-riscv-acpi-table-upgrade-v1-1-45902d2dedf9@iscas.ac.cn
Signed-off-by: Paul Walmsley <pjw@kernel.org>
|
|
Firmware-preferred reset and EFI capsule update support requires reset
via EFI runtime services rather than direction M-mode firmware invocation
via SBI. Unlike poweroff, restart mechanism is directly controlled from
machine_restart function though.
Prefer the EFI runtime ResetSystem() service for restart when UEFI
runtime services are available.
Signed-off-by: Atish Patra <atishp@meta.com>
Reviewed-by: Sunil V L <sunilvl@oss.qualcomm.com>
Link: https://patch.msgid.link/20260615-efi_reset_shutdown-v1-2-9414edcbbab0@meta.com
Signed-off-by: Paul Walmsley <pjw@kernel.org>
|
|
When booted via UEFI with runtime services enabled, EFI Reset Shutdown is
the firmware-preferred shutdown path: it lets firmware run its own
shutdown hooks which may invoke SBI SRST extension internally. However,
RISC-V always powers off via the SBI SRST extension today and EFI runtime
path is never used even when firmware provides it.
Enable the poweroff via EFI by overriding efi_poweroff_required()
Signed-off-by: Atish Patra <atishp@meta.com>
Reviewed-by: Sunil V L <sunilvl@oss.qualcomm.com>
Link: https://patch.msgid.link/20260615-efi_reset_shutdown-v1-1-9414edcbbab0@meta.com
Signed-off-by: Paul Walmsley <pjw@kernel.org>
|
|
Ivan has been continuously active in DPLL development and discussion
since April 2025. His contributions cover the DPLL core and API,
netlink, bindings, ICE/SyncE integration, and the ZL3073x driver.
He also regularly reviews and tests DPLL patches from other contributors
and already maintains the Microchip ZL3073x driver. Add him as a reviewer
to reflect his ongoing involvement across the subsystem.
Signed-off-by: Jiri Pirko <jiri@nvidia.com>
Acked-by: Vadim Fedorenko <vadim.fedorenko@linux.dev>
Acked-by: Arkadiusz Kubalewski <arkadiusz.kubalewski@intel.com>
Link: https://patch.msgid.link/20260806094432.163833-1-jiri@resnulli.us
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|