From nobody Mon Sep 28 17:48:47 2026 Received: from mail-pz2-f0.google.com (mail-pz2-f0.google.com [74.125.228.0]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 7EE76431A41 for ; Wed, 19 Aug 2026 09:15:13 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.228.0 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787130927; cv=none; b=IMBKEJMJZMy4aSdF2e2vHonausv7k22am6dX8N+ByUDcLQbJ8xBna0qr2RaRmuZDtZ4MPJlVuuzysWi8U44ho8+CfIkFPKlMODgKN0tEJqOwBA6mOAPWAA2Md3MvpQYl4s85enJGNuXExy7krVonNc1lgycuSyIxAqUC2fxSDEc= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787130927; c=relaxed/simple; bh=gnF986okVh0Ca6Epmk06H19lpWBkHyfiiPqLqeiSQbY=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=Pzz9Iq0iULLorvrMF0OSyDFu/iyKxo84bNf6wC5zr+YDJHlO6+1rPXUb1DVaXZDryUqmo1ZgYXpwwPf0AHoAsY9eztgotynewE0NAtTJnTC1poBizL9xObumkb5etdOLYynmHDe25QtvWpYad7cvZ7Uv3To9k4ayURAOEWGjI6k= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=ZUN9ZMay; arc=none smtp.client-ip=74.125.228.0 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="ZUN9ZMay" Received: by mail-pz2-f0.google.com with SMTP id 41be03b00d2f7-ca92f36c4acso191344a12.1 for ; Wed, 19 Aug 2026 02:15:13 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1787130911; x=1787735711; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=JPqYiA3RKzrqvqwgzAS3BUvxOhZnKEBVLQLALJPs0HI=; b=ZUN9ZMayUyVOQvgZBgnsAYuLSKv6hPOcoUNC6m1uEm4d7sTG85DUdXZOpk2bcFVz1g F4S8/U6ghm+Zx4uvhClxL1ggbSn8Shm1LzjmM41tVxdqtGLHmTDtkFTraYgC/eUcmGkb iiddZJd4sFP4DBytZDy8qm1V4iI+EC5k3IP4+8T7cj1Mb81g/uoX5iGkl1GeUM2NBNjs y+0bSzGVQkSDf98cRv7blg+qykvrKxIzifeC6zzDVEg6oEXPa7jPU9p3RjHVK2qCffQx NX8SE3Mob1NvvYg67a3L4Hl5GZd9cQUAYitLBIl2rCoUZ9XcIMvBVoogCVxPTwBNsM2Y wRqA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1787130911; x=1787735711; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=JPqYiA3RKzrqvqwgzAS3BUvxOhZnKEBVLQLALJPs0HI=; b=fpljbCZgs0HV9Z43H9GorknMIcCLFwgAau0R5h1QjgDVrCIFgAfzaktKAO4d7BYcS3 mIOChL+mRvk+lCIe8rKjmrSmsg3f5UfjJXYEapMk90PPwbYtTFw9fxRwGUCKka4VXRHI PU8jhNLra7vaCdu7Z8nidvBN8JwlmgufnKCN6DoLR09p18ej+Pnt110qq94KoOGPZqSJ Nb7kUYcTp+b6KgOVGOdaukzaMnxSIQmzSm0kwgFAOmRd6F4NFowOiWM0gpCuiXOnQ5Og TD1stBnY6eR2TGdjwVMtkzACINfDrQdJ9ic2k1kdysUy5YbWlbVSKpwaZ5xBKp0KGFDM DyJA== X-Forwarded-Encrypted: i=1; AHgh+RpEbGj7nsXNrw228X/Yn0cwbESH/psfRhUi+pnZEhu8LN4HZFyucY18UyDFfIVaBeZbhIsaZEudhuEPJWE=@vger.kernel.org X-Gm-Message-State: AOJu0YwjrLUNYLS7wEY9REt0aj3ViLQRDNPYqldC7O5sH6YpXMyIpvZS FVFEEnuZg+ZfcrvEKFqvgooPX7DsQHfYjo5bwdzbNOjRqva0F3djW1Wp X-Gm-Gg: AR+sD13ojPEzqL6gzCzp4eRazBw/13LIfqjTKuKIBcIxJC4/DeW2PvG8hD/FwMZSSJP +VlX6zQA9n1mkchIgAlZRzdQkR/iYl1wV6V6tY01DPBuuYEDHnLaPvyB51A3jwQYkjZ5IbyYOYi a9uOvsbzPrGgTFBUI4vEp0xQPeRX5NT83uRoKmGCNR/ivZfhic68KSDAV8ZZMfd+fMFiqUOBah9 nqBqCKqaJiivtvlLhC2Hf4oygV68Ala06r/z1VpAl4Dmddnhv3v46JjYFBjIcpQhv9XDz4FpwrS 3P3rIi/5PXJcCVzO4RfI15vOLIesoENmTbncVtHRXAjwm2M31KbxpkTE2BVGYNEv+6QT5Y2w+t6 PagK6oO81I3/gfxqjVj6+tHuI9QC4ttjQd3h24az1nKPBhVaw44cBy677mHd/pvnEjAStVUn7J7 FLsSBRffbCRBfEHdsOeA8DvCHch9BozKBhcTkfgiIdIuT/OkQ/edOLQCd30vHUD8yiTMpL20iuT mkO X-Received: by 2002:a05:6a00:368f:b0:84e:f90e:492f with SMTP id d2e1a72fcca58-851d382f530mr5371908b3a.7.1787130911266; Wed, 19 Aug 2026 02:15:11 -0700 (PDT) Received: from otkcomputer.localdomain ([111.68.15.147]) by smtp.gmail.com with ESMTPSA id d2e1a72fcca58-851d36a2ef0sm393956b3a.61.2026.08.19.02.15.06 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Wed, 19 Aug 2026 02:15:10 -0700 (PDT) From: syhuang.nju@gmail.com To: brauner@kernel.org Cc: aliceryhl@google.com, dianders@chromium.org, gary@garyguo.net, jack@suse.cz, linux-fsdevel@vger.kernel.org, linux-kernel@vger.kernel.org, linux.amoon@gmail.com, nathan@kernel.org, nsc@kernel.org, ojeda@kernel.org, syhuang.nju@gmail.com, tglx@kernel.org, thomas.weissschuh@linutronix.de, viro@zeniv.linux.org.uk Subject: [RFC PATCH v2] eventpoll: add lockless fast-path for eventpoll Date: Wed, 19 Aug 2026 17:15:01 +0800 Message-ID: <20260819091501.3935-1-syhuang.nju@gmail.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260812-zitat-neueinstellung-ruckartig-727e55f45cc1@brauner> References: <20260812-zitat-neueinstellung-ruckartig-727e55f45cc1@brauner> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Siyuan Huang > I think this is misguided. Not just the special-casing of eventfd but > also letting file operations opt-in how locking is done in another > subystem. That's almost guaranteed to wreak havoc... and sashiko is > already proving this. Thanks for the feedback. I agree that the fast-path logic should not be connected to other subsystems. I have reworked the proposal so that the optimization is entirely local to eventpoll. I removed the special-casing of eventfd and the opt-in bit in FOP. The new version tracks two bits in each epitem: - FAST marks an item that is eligible for the lockless callback path. Nested epoll and interests using EPOLLEXCLUSIVE, EPOLLET, EPOLLONESHOT, or EPOLLWAKEUP do not use the fast path. - ACCOUNTED means that the ready-processing path already owns the item: it is on rdllist, in the current scan batch, or queued on ovflist, the same as the 'notified' before. As for the question sashiko raised about the fatal branch, I do not modify the flag in the branch any more. Instead, I deliver the wakeup directly to avoid list traversal race. --- fs/eventpoll.c | 101 ++++++++++++++++++++++++++++++++++++++++++++++++- 1 file changed, 99 insertions(+), 2 deletions(-) diff --git a/fs/eventpoll.c b/fs/eventpoll.c index eed8cecd94e3..2c29f5d0039f 100644 --- a/fs/eventpoll.c +++ b/fs/eventpoll.c @@ -243,6 +243,14 @@ struct eppoll_entry { wait_queue_head_t *whead; }; =20 +/* + * FAST is set for eligible items at insertion. ACCOUNTED tracks FAST items + * while the ready path owns them and is cleared immediately before re-pol= l. + */ +enum { + EP_STATE_ACCOUNTED =3D BIT(0), + EP_STATE_FAST =3D BIT(1), +}; /* * Each file descriptor added to the eventpoll interface will * have an entry of this type linked to the "rbr" RB tree. @@ -271,6 +279,9 @@ struct epitem { /* The file descriptor information this item refers to */ struct epoll_key ffd; =20 + /* Atomic state bits, see the EP_STATE_* enum above. */ + atomic_t state; + /* List containing poll wait queues */ struct eppoll_entry *pwqlist; =20 @@ -620,6 +631,65 @@ static inline bool ep_events_available(struct eventpol= l *ep) read_seqcount_retry(&ep->seq, seq); } =20 +/* + * Let A be ACCOUNTED and B be file readiness. Fully ordered atomic RMWs + * pair W(A=3D0) before R(B) with W(B) before R(A), so a scanner cannot mi= ss + * B while its callback skips an accounted item. ep_done_scan() and the + * fatal-signal handoff pass queued events between exclusive waiters. + * poll_wait users, private modes and special callbacks take the slow path. + */ +static inline bool ep_state_test(const struct epitem *epi, unsigned int bi= t) +{ + return atomic_read(&epi->state) & bit; +} + +static inline bool ep_ready_fast_enabled(const struct epitem *epi) +{ + return ep_state_test(epi, EP_STATE_FAST); +} + +static inline void ep_set_ready_accounted(struct epitem *epi) +{ + if (ep_ready_fast_enabled(epi)) + atomic_or(EP_STATE_ACCOUNTED, &epi->state); +} + +static inline void ep_prepare_repoll(struct epitem *epi) +{ + if (ep_ready_fast_enabled(epi)) + atomic_fetch_andnot(EP_STATE_ACCOUNTED, &epi->state); +} + +static bool ep_callback_can_skip(struct epitem *epi, __poll_t pollflags) +{ + __poll_t events; + + events =3D READ_ONCE(epi->event.events); +=09 + /* Only insertion-time candidates may use the lockless path. */ + if (!ep_ready_fast_enabled(epi)) + return false; +=09 + /* ep_modify() can select a private mode that requires the slow path. */ + if (events & EP_PRIVATE_BITS) + return false; + + /* poll_wait users must be woken by this callback. */ + if (waitqueue_active(&epi->ep->poll_wait)) + return false; + + /* POLLFREE tears down wait entries; URING_WAKE must propagate. */ + if (pollflags & (POLLFREE | EPOLL_URING_WAKE)) + return false; + + /* Unmatched wake keys retain the slow-path callback semantics. */ + if (pollflags && !(pollflags & events)) + return false; + + /* Fully ordered RMW pairs the B update with the scanner's A clear. */ + return atomic_fetch_or(0, &epi->state) & EP_STATE_ACCOUNTED; +} + #ifdef CONFIG_NET_RX_BUSY_POLL /** * busy_loop_ep_timeout - check if busy poll has timed out. The timeout va= lue @@ -1007,6 +1077,7 @@ static void ep_done_scan(struct eventpoll *ep, * reverses the iteration order into FIFO. */ list_add(&epi->rdllink, &ep->rdllist); + ep_set_ready_accounted(epi); ep_pm_stay_awake(epi); } } @@ -1303,8 +1374,10 @@ static __poll_t __ep_eventpoll_poll(struct file *fil= e, poll_table *wait, int dep mutex_lock_nested(&ep->mtx, depth); ep_start_scan(ep, &scan_batch); list_for_each_entry_safe(epi, tmp, &scan_batch, rdllink) { + ep_prepare_repoll(epi); if (ep_item_poll(epi, &pt, depth + 1)) { res =3D EPOLLIN | EPOLLRDNORM; + ep_set_ready_accounted(epi); break; } else { /* @@ -1497,6 +1570,9 @@ static int ep_poll_callback(wait_queue_entry_t *wait,= unsigned mode, int sync, v unsigned long flags; int ewake =3D 0; =20 + if (ep_callback_can_skip(epi, pollflags)) + return 1; + spin_lock_irqsave(&ep->lock, flags); =20 ep_set_busy_poll_napi_id(epi); @@ -1529,11 +1605,13 @@ static int ep_poll_callback(wait_queue_entry_t *wai= t, unsigned mode, int sync, v if (!epi_on_ovflist(epi)) { epi->ovflist_next =3D READ_ONCE(ep->ovflist); WRITE_ONCE(ep->ovflist, epi); + ep_set_ready_accounted(epi); ep_pm_stay_awake_rcu(epi); } } else if (!ep_is_linked(epi)) { /* In the usual case, add event to ready list. */ list_add_tail(&epi->rdllink, &ep->rdllist); + ep_set_ready_accounted(epi); ep_pm_stay_awake_rcu(epi); } =20 @@ -1912,6 +1990,13 @@ static int ep_insert(struct ep_ctl_ctx *ctx, struct = eventpoll *ep, if (IS_ERR(epi)) return PTR_ERR(epi); =20 + /* + * Lockless candidates are plain level-triggered interests; + * private modes and nested epoll change the callback contract. + */ + if (!tep && !(event->events & EP_PRIVATE_BITS)) + atomic_set(&epi->state, EP_STATE_FAST); + error =3D ep_register_epitem(ctx, ep, epi, tep, full_check); if (error) return error; @@ -1956,6 +2041,7 @@ static int ep_insert(struct ep_ctl_ctx *ctx, struct e= ventpoll *ep, =20 if (revents && !ep_is_linked(epi)) { list_add_tail(&epi->rdllink, &ep->rdllist); + ep_set_ready_accounted(epi); ep_pm_stay_awake(epi); =20 if (waitqueue_active(&ep->wq)) @@ -1992,7 +2078,8 @@ static int ep_modify(struct eventpoll *ep, struct epi= tem *epi, * otherwise we might miss an event that happens between the * f_op->poll() call and the new event set registering. */ - epi->event.events =3D event->events; /* need barrier below */ + WRITE_ONCE(epi->event.events, event->events); /* need barrier below */ epi->event.data =3D event->data; /* protected by mtx */ if (epi->event.events & EPOLLWAKEUP) { if (!ep_has_wakeup_source(epi)) @@ -2031,6 +2118,7 @@ static int ep_modify(struct eventpoll *ep, struct epi= tem *epi, spin_lock_irq(&ep->lock); if (!ep_is_linked(epi)) { list_add_tail(&epi->rdllink, &ep->rdllist); + ep_set_ready_accounted(epi); ep_pm_stay_awake(epi); =20 /* Notify waiting tasks that events are available */ @@ -2084,6 +2172,7 @@ static int ep_deliver_event(struct eventpoll *ep, str= uct epitem *epi, __pm_relax(ws); } =20 + ep_prepare_repoll(epi); list_del_init(&epi->rdllink); =20 /* @@ -2104,13 +2193,15 @@ static int ep_deliver_event(struct eventpoll *ep, s= truct epitem *epi, * attempt. */ list_add(&epi->rdllink, scan_batch); + ep_set_ready_accounted(epi); ep_pm_stay_awake(epi); return -EFAULT; } *uevents =3D next; =20 if (epi->event.events & EPOLLONESHOT) { - epi->event.events &=3D EP_PRIVATE_BITS; + WRITE_ONCE(epi->event.events, + READ_ONCE(epi->event.events) & EP_PRIVATE_BITS); } else if (!(epi->event.events & EPOLLET)) { /* * Level-triggered: re-queue so the next epoll_wait() @@ -2120,6 +2211,7 @@ static int ep_deliver_event(struct eventpoll *ep, str= uct epitem *epi, * during scans. */ list_add_tail(&epi->rdllink, &ep->rdllist); + ep_set_ready_accounted(epi); ep_pm_stay_awake(epi); } return 1; @@ -2288,6 +2380,11 @@ static int ep_poll(struct eventpoll *ep, struct epol= l_event __user *events, while (1) { if (eavail) { res =3D ep_try_send_events(ep, events, maxevents); + if (res =3D=3D -EINTR) { + spin_lock_irq(&ep->lock); + wake_up(&ep->wq); + spin_unlock_irq(&ep->lock); + } if (res) return res; } --=20 2.43.0