[PATCH] eventfs: Add warning for out of bounds pos in __eventfs_iterate()

Steven Rostedt posted 1 patch 1 month, 2 weeks ago
fs/tracefs/event_inode.c | 4 ++++
1 file changed, 4 insertions(+)
[PATCH] eventfs: Add warning for out of bounds pos in __eventfs_iterate()
Posted by Steven Rostedt 1 month, 2 weeks ago
From: Steven Rostedt <rostedt@goodmis.org>

Sashiko has complained about out of bounds issues if ctx->pos isn't what
is expected in __eventfs_iterate()[1]. This would be an issue if the logic
that calls __eventfs_iterate() didn't already prevent the code from going
out of bounds.

The issue Sashiko brings up is if a user uses lseek64() to put in a
position like 0x100000000 which will overflow the integer used to iterate
the files. This should never be an issue because both tracefs and eventfs
uses the default "maxbytes" for its superblock "s_maxbytes" field which is
defined as:

  fs/super.c:     s->s_maxbytes = MAX_NON_LFS;
  include/linux/fs.h:#define      MAX_NON_LFS     ((1UL<<31) - 1)

Where MAX_NON_LFS turns into 0x7fffffff.

Testing this with code to try to pass 0x100000000 to lseek64() to a
eventfs directory returns -EINVAL.

But relying on logic for the integrity of a function is not very robust.
Add a WARN_ON_ONCE() in case the ctx->pos is out of the expected range.

[1] https://sashiko.dev/#/patchset/20260810160708.3460a2fd%40gandalf.local.home

Signed-off-by: Steven Rostedt <rostedt@goodmis.org>
---
 fs/tracefs/event_inode.c | 4 ++++
 1 file changed, 4 insertions(+)

diff --git a/fs/tracefs/event_inode.c b/fs/tracefs/event_inode.c
index be21b8510c14..ee5d03e403ec 100644
--- a/fs/tracefs/event_inode.c
+++ b/fs/tracefs/event_inode.c
@@ -696,6 +696,10 @@ static int __eventfs_iterate(struct file *file, struct dir_context *ctx, bool ro
 	if (!(ti->flags & TRACEFS_EVENT_INODE))
 		return -EINVAL;
 
+	/* Logic should prevent ctx->pos from going out of range */
+	if (WARN_ON_ONCE(ctx->pos < 2 || ctx->pos > 0x7fffffffULL))
+		return -EINVAL;
+
 	c = ctx->pos - 2;
 
 	guard(srcu)(&eventfs_srcu);
-- 
2.53.0
Re: [PATCH] eventfs: Add warning for out of bounds pos in __eventfs_iterate()
Posted by Steven Rostedt 1 month, 2 weeks ago
Replying to Sashiko: https://sashiko.dev/#/patchset/20260810175928.5f4d9d5c%40gandalf.local.home

> --- Patch [1]: [PATCH] eventfs: Add warning for out of bounds pos in __eventfs_iterate() ---
> commit 2f3bdca57e4eb4f6f462c12ceab80527f383bb86
> Author: Steven Rostedt <rostedt@goodmis.org>
> 
> eventfs: Add warning for out of bounds pos in __eventfs_iterate()
> 
> This commit adds a WARN_ON_ONCE to eventfs_iterate() to catch out of bounds
> positions. It ensures that ctx->pos remains within the expected range, 
> guarding against potential overflow issues from 64-bit seeks.
> 
> Link: https://sashiko.dev/#/patchset/20260810160708.3460a2fd%40gandalf.local.home
> 
> > diff --git a/fs/tracefs/event_inode.c b/fs/tracefs/event_inode.c
> > index a524584353279..604ba3e841d24 100644
> > --- a/fs/tracefs/event_inode.c
> > +++ b/fs/tracefs/event_inode.c
> > @@ -594,6 +594,10 @@ static int eventfs_iterate(struct file *file, struct dir_context *ctx)
> >  	if (!(ti->flags & TRACEFS_EVENT_INODE))
> >  		return -EINVAL;
> >  
> > +	/* Logic should prevent ctx->pos from going out of range */
> > +	if (WARN_ON_ONCE(ctx->pos < 2 || ctx->pos > 0x7fffffffULL))
> > +		return -EINVAL;
> > +
> >  	c = ctx->pos - 2;
> >  
> >  	guard(srcu)(&eventfs_srcu);
> 
> [ ... ]
> 
> >  	list_for_each_entry_srcu(ei_child, &ei->children, list,
> >  				 srcu_read_lock_held(&eventfs_srcu)) {
> 
> [ ... ]
> 
> >  		smp_rmb();
> >  		if (ei->is_freed)
> >  			return -EINVAL;
> >  
> >  		if (c > 0) {
> >  			c--;
> >  			continue;
> >  		}
> 
> This is a pre-existing issue, but does the skip logic here using the
> absolute index c risk omitting valid directory entries during concurrent
> modifications?
> 
> If a user calls getdents on an eventfs directory and it returns partially
> (e.g., due to a full buffer), and a child entry is concurrently removed from
> ei->children via list_del_rcu(), the list shrinks.
> 
> When getdents resumes on the next syscall, eventfs_iterate() recalculates
> c = ctx->pos - 2. Since the list physically shrank, using this absolute
> index might skip too many elements in the now-shrunken list, bypassing valid
> entries that have shifted forward.
> 
> Is there a mechanism preventing this race between getdents and concurrent
> removals?

No and we do not care. The tracefs / eventfs file system is only modified
when new events are added or removed. During those transactions, reading
the directory has an undefined behavior. Always has, and always will. If
events are dropped because the admin is both reading the directories as
they add or remove events, then it is expected that the result may be
garbage.

-- Steve