From nobody Sat Jul 25 21:22:49 2026 Received: from mail-pj1-f48.google.com (mail-pj1-f48.google.com [209.85.216.48]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id D985F28B4E2 for ; Mon, 13 Jul 2026 16:06:02 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.216.48 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1783958765; cv=none; b=J44MwezPR0EUSfCEHDev0+4B/m4CJnX/03FZ3HJtIqX70I6GF/fYcDtPsoaJto9CmRtEWEibS2YL+Ph3+wev0cepx5zi1MATQ0eTgB91JGUfgWPrI/brvW0mjoew+fg/PSlccVLTTezaOho+nPQyPCxitdS3nHYCS0Ft6AD3/M0= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1783958765; c=relaxed/simple; bh=Q8FjjEFMAnuE8UmoPXbF5lB6kQ+89kwNqcgpa2DoFp8=; h=From:To:Cc:Subject:Date:Message-ID:MIME-Version; b=WROnQI5eb59L0v5LPGMf4myOXzR4D/LrL5KaA5QUm3lvNR/ti5m0b4UHWUNgZdKrLDARAvLWfMHqhSZBE8nfH6U0a5Z8ZkLAUb4ehZVYidLtmW1Hbbf1J7uQK3+AdDxujFxR9n53JvDuTlpjTyM2vYJm6G9abWnK0ADzI0vYdMA= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=DOMmaeyw; arc=none smtp.client-ip=209.85.216.48 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="DOMmaeyw" Received: by mail-pj1-f48.google.com with SMTP id 98e67ed59e1d1-37e0a189b0bso78710a91.1 for ; Mon, 13 Jul 2026 09:06:02 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1783958762; x=1784563562; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:message-id:date:subject:cc :to:from:from:to:cc:subject:date:message-id:reply-to:content-type; bh=4XAW84Oy7NZ27UgySG9rIfuC0UVJRqbx2WnYyVZK3ys=; b=DOMmaeyw/C9vWtp2uDK2rweNEJnpLMguTCx+6O+aHJdoInTzLMgfjN+MobHmyD/lEO 1CXfwDkRkB91875/q2I6nF7NhItCplgOTHNm0Yc3LlggrTBzQSaxU9ytY8L0AVN6DfEb BFhMLjon5eCNpbi96+QldC7xDX1xMfytY9CZkTRaBKlFfCvjFo7Uduib+9qY4pVwbzrS gOK/faMDsbNw/E1MaS2bRQ4x2+Ry4Y75RBklGkri5jKeEmIxykEiXU6wx9cDLVYJmYsu gWGgjtQ0owjUQoRG9z/4E/PSdz6qUrSK3qlDTvIEKWRVk4YWz4UWwJ4sFq/zIANhDg/9 CUyQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1783958762; x=1784563562; h=content-transfer-encoding:mime-version:message-id:date:subject:cc :to:from:x-gm-gg:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=4XAW84Oy7NZ27UgySG9rIfuC0UVJRqbx2WnYyVZK3ys=; b=WvRcK6HRXUaJ6JFLFYCmNHOZnK7KCJMhK/40l4Dps+l4BtI3R1ohbsMuigITF3FPB8 h8aM9ioCRkpD32Bjm7zxCRtWn/LiG4yu+2XTwi2N7Jw6VcDJoHM9HSvpvsnoANYjyvIm h2mK5/ln4NVdPEfxM5bGCDPCTnXUsEztB5TsIlm5sG+s+DnEEGjVYfJACAOevyxgH/Wn iEQO+cUTe/fWbwgPPwSUbHL8y6c1s+lknvcmhFo5apjxyzYD6vYwUtToawyecEXSZoYp lm15w8xkFhEDQ5fE+GhTjT1JrL+7YfsA4v4TVyInz3wC9OD6BMZIPtm780ywgUjGg+T9 DOcA== X-Gm-Message-State: AOJu0YyfHIh7qQ3dhTvQEKabUCQshhWxAW2r6lzjo5rSYOqBVdUewjk1 U0My4WkpU1DEiyKis9PgcJy8Y49KLTdOpEGJ8WEx9ugTwmKW5/6ntQc9RzX04A== X-Gm-Gg: AfdE7cmwsnyNbbpW17sWdqLcuL5x9IhrG0oudM1MKr9rZor7d90SP0HNsn17ZPC8C3D Wpq7HwH/fLEcPprIDPGMzj5GR01xiKRSmKp2Yik+paLZgnC7Qo0dULtRxUuXi/DRul9bcgRhk/u 3fNUh3bUS4qIfMh+mqe74z5rurUhn9S2HGovg2EacoJGBMJ5361Yx/7yCPQEIyXXkdPEehHNMMy qldVSnxNSgNIRnl1f7YMjk93lBJx5wM0ZJR5cEMY6gIobAdZwr5YTZu4BILLkHmmB3uzFgwAENX 8R2YArWmXOqEVzy1mACRZKJz4ImecscP9+NufrgVlmopV8fLbp8l4drSL5Mk39Y/Nx4i0yEJKn0 7ZJXL7Ggli9mYThzz/stLENyQxhcT5Yw7wMnYI5C2U2XarI5GRNTAB40JQSfOe+bgTjQs/CSvrm LabDLJ2RGxLK+X1rowRRUzjb0+pwTbdoAYfb0mzwBNBHBHtwFdUMqUuusmlAA0g2ok0SRz/9NMe ceIE2wRNFJcrikmkoFrA8QCGluQmNooNYBvhvE= X-Received: by 2002:a17:90b:2f8b:b0:382:3dcc:1487 with SMTP id 98e67ed59e1d1-38dc7b3f5demr9038598a91.25.1783958761636; Mon, 13 Jul 2026 09:06:01 -0700 (PDT) Received: from daehojeong-desktop.mtv.corp.google.com ([2a00:79e0:2e7c:8:1f9:fdeb:bef5:6a1b]) by smtp.gmail.com with ESMTPSA id 5a478bee46e88-31198cb2b99sm47819071eec.26.2026.07.13.09.06.00 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Mon, 13 Jul 2026 09:06:00 -0700 (PDT) From: Daeho Jeong To: linux-kernel@vger.kernel.org, linux-f2fs-devel@lists.sourceforge.net, kernel-team@android.com Cc: Daeho Jeong Subject: [PATCH v4] f2fs: support dynamic reserve/release for device aliasing Date: Mon, 13 Jul 2026 09:05:56 -0700 Message-ID: <20260713160556.3988119-1-daeho43@gmail.com> X-Mailer: git-send-email 2.55.0.795.g602f6c329a-goog Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Daeho Jeong This patch adds a dynamic management feature to the existing device aliasing functionality. It allows users to dynamically reserve or release specific devices from the filesystem's free pool at runtime through new ioctls. To support this, three new ioctls are introduced: - F2FS_IOC_RESERVE_DEV_ALIAS: This reclaims the space occupied by a device aliasing file. It first performs a capacity check, resets GC victim information for the target range, marks the segments as in-use to prevent new allocations, and then triggers GC to migrate existing valid data out of the range. Finally, it reserves these blocks in the SIT to effectively exclude the device from the usable capacity. - F2FS_IOC_RELEASE_DEV_ALIAS: This releases the reserved space of a previously reserved device aliasing file. It truncates the blocks associated with the file, which makes them available for general filesystem allocation again. - F2FS_IOC_GET_DEV_ALIAS_STATUS: This retrieves the current aliasing status of a device aliasing file, returning whether the file is released (inactive alias) or reserved (active alias, with blocks fully allocated on the device). Signed-off-by: Daeho Jeong --- v4: renamed interfaces. fixed race conditions between checkpoint=3Ddisable mount and ioctls. refactored segment reservation part. modified lock usage. v3: add CAP_SYS_ADMIN and checkpoint=3Ddisabled check. remove a f2fs specific flag exposed with getflags. v2: prevent operations during checkpoint=3Ddisabled. --- Documentation/filesystems/f2fs.rst | 35 ++++ fs/f2fs/f2fs.h | 9 +- fs/f2fs/file.c | 289 ++++++++++++++++++++++++++++- fs/f2fs/gc.c | 30 +-- fs/f2fs/namei.c | 14 ++ fs/f2fs/segment.c | 180 +++++++++++++----- fs/f2fs/segment.h | 26 +++ fs/f2fs/super.c | 34 ++++ include/uapi/linux/f2fs.h | 7 + 9 files changed, 562 insertions(+), 62 deletions(-) diff --git a/Documentation/filesystems/f2fs.rst b/Documentation/filesystems= /f2fs.rst index 8c4a14ae444f..1a5fd4afe609 100644 --- a/Documentation/filesystems/f2fs.rst +++ b/Documentation/filesystems/f2fs.rst @@ -1045,6 +1045,41 @@ So, the key idea is, user can do any file operations= on /dev/vdc, and reclaim the space after the use, while the space is counted as /data. That doesn't require modifying partition size and filesystem format. =20 +Dynamic Device Aliasing Management +---------------------------------- + +In addition to static device aliasing by deleting the aliasing file, F2FS +supports dynamic management of device aliasing. This mechanism allows the = system +to dynamically transition partition ownership between F2FS userdata and ex= ternal +entities (e.g., zRAM, raw partition) based on system requirements without +deleting the master aliasing file or requiring unmount/remount. + +The master aliasing file is created during the initial format of the file = system +and remains as a persistent control entity (ioctl gateway) in the root dir= ectory. + +- Partition Reservation (In-service to Aliased) + When a specific partition needs to be dedicated to external services (e.= g., zRAM), + a user can reserve the device alias range via ioctl. The kernel resets G= C victim + information for the target range, marks segments as in-use to prevent new + allocations, and triggers forced GC to migrate existing valid data out o= f the + range. Finally, it reserves these blocks in the SIT to effectively exclu= de the + device from the usable capacity. + +- Partition Release (Aliased to In-service) + When external usage concludes, the space is reclaimed not by deleting th= e file, + but through the release ioctl. The kernel truncates blocks associated wi= th + the file, releasing them back to general filesystem allocation. + +.. code-block:: + + # f2fs_io dev_alias release /mnt/f2fs/vdc.file + # df -h + /dev/vdb 64G 753M 64G 2% /mnt/f2fs + + # f2fs_io dev_alias reserve /mnt/f2fs/vdc.file + # df -h + /dev/vdb 64G 33G 32G 52% /mnt/f2fs + Per-file Read-Only Large Folio Support -------------------------------------- =20 diff --git a/fs/f2fs/f2fs.h b/fs/f2fs/f2fs.h index f1774d4e18d2..0f9b8b66cef9 100644 --- a/fs/f2fs/f2fs.h +++ b/fs/f2fs/f2fs.h @@ -1404,6 +1404,8 @@ struct f2fs_dev_info { unsigned int total_segments; block_t start_blk; block_t end_blk; + bool has_alias; + bool is_reserving; #ifdef CONFIG_BLK_DEV_ZONED unsigned int nr_blkz; /* Total number of zones */ unsigned long *blkz_seq; /* Bitmap indicating sequential zones */ @@ -4009,7 +4011,10 @@ int f2fs_create_flush_cmd_control(struct f2fs_sb_inf= o *sbi); int f2fs_flush_device_cache(struct f2fs_sb_info *sbi); void f2fs_destroy_flush_cmd_control(struct f2fs_sb_info *sbi, bool free); void f2fs_invalidate_blocks(struct f2fs_sb_info *sbi, block_t addr, - unsigned int len); + unsigned int len); +void f2fs_reserve_device_alias(struct f2fs_sb_info *sbi, block_t addr, + unsigned int len); + bool f2fs_is_checkpointed_data(struct f2fs_sb_info *sbi, block_t blkaddr); int f2fs_start_discard_thread(struct f2fs_sb_info *sbi); void f2fs_drop_discard_cmd(struct f2fs_sb_info *sbi); @@ -4231,6 +4236,8 @@ void f2fs_build_gc_manager(struct f2fs_sb_info *sbi); int f2fs_gc_range(struct f2fs_sb_info *sbi, unsigned int start_seg, unsigned int end_seg, bool dry_run, unsigned int dry_run_sections); +void f2fs_reset_gc_victim_resource(struct f2fs_sb_info *sbi, + unsigned int start, unsigned int end); int f2fs_resize_fs(struct file *filp, __u64 block_count); int __init f2fs_create_garbage_collection_cache(void); void f2fs_destroy_garbage_collection_cache(void); diff --git a/fs/f2fs/file.c b/fs/f2fs/file.c index 4b52c56d71f0..9077c091c2d6 100644 --- a/fs/f2fs/file.c +++ b/fs/f2fs/file.c @@ -813,13 +813,19 @@ int f2fs_do_truncate_blocks(struct inode *inode, u64 = from, bool lock) =20 if (IS_DEVICE_ALIASING(inode)) { struct extent_tree *et =3D F2FS_I(inode)->extent_tree[EX_READ]; - struct extent_info ei =3D et->largest; + struct extent_info ei; + + read_lock(&et->lock); + ei =3D et->largest; + read_unlock(&et->lock); =20 f2fs_invalidate_blocks(sbi, ei.blk, ei.len); =20 dec_valid_block_count(sbi, inode, ei.len); f2fs_update_time(sbi, REQ_TIME); =20 + f2fs_drop_extent_tree(inode); + f2fs_folio_put(ifolio, true); goto out; } @@ -1100,8 +1106,9 @@ int f2fs_setattr(struct mnt_idmap *idmap, struct dent= ry *dentry, if ((attr->ia_valid & ATTR_SIZE)) { if (mapping_large_folio_support(inode->i_mapping)) return -EOPNOTSUPP; - if (!f2fs_is_compress_backend_ready(inode) || - IS_DEVICE_ALIASING(inode)) + if (IS_DEVICE_ALIASING(inode)) + return -EPERM; + if (!f2fs_is_compress_backend_ready(inode)) return -EOPNOTSUPP; if (is_inode_flag_set(inode, FI_COMPRESS_RELEASED) && !IS_ALIGNED(attr->ia_size, @@ -2130,6 +2137,9 @@ static int f2fs_setflags_common(struct inode *inode, = u32 iflags, u32 mask) if (IS_NOQUOTA(inode)) return -EPERM; =20 + if (IS_DEVICE_ALIASING(inode)) + return -EPERM; + if ((iflags ^ masked_flags) & F2FS_CASEFOLD_FL) { if (!f2fs_sb_has_casefold(F2FS_I_SB(inode))) return -EOPNOTSUPP; @@ -2678,6 +2688,17 @@ static int f2fs_ioc_get_encryption_policy(struct fil= e *filp, unsigned long arg) return fscrypt_ioctl_get_policy(filp, (void __user *)arg); } =20 +static int f2fs_ioc_get_dev_alias_status(struct file *filp, unsigned long = arg) +{ + struct inode *inode =3D file_inode(filp); + + if (!IS_DEVICE_ALIASING(inode)) + return -EINVAL; + + return put_user(F2FS_HAS_BLOCKS(inode) ? F2FS_DEV_ALIAS_STATUS_RESERVED : + F2FS_DEV_ALIAS_STATUS_RELEASED, (u32 __user *)arg); +} + static int f2fs_ioc_get_encryption_pwsalt(struct file *filp, unsigned long= arg) { struct inode *inode =3D file_inode(filp); @@ -3616,6 +3637,259 @@ static int f2fs_ioc_get_dev_alias_file(struct file = *filp, unsigned long arg) (u32 __user *)arg); } =20 +static int f2fs_ioc_reserve_dev_alias(struct file *filp) +{ + struct inode *inode =3D file_inode(filp); + struct f2fs_sb_info *sbi =3D F2FS_I_SB(inode); + struct extent_tree *et =3D F2FS_I(inode)->extent_tree[EX_READ]; + struct extent_info ei; + struct cp_control cpc =3D { CP_SYNC, 0, 0, 0 }; + struct f2fs_lock_context lc, glc; + blkcnt_t count; + unsigned int start, end, segno; + int type, i, err; + + if (!capable(CAP_SYS_ADMIN)) + return -EPERM; + + if (unlikely(is_sbi_flag_set(sbi, SBI_CP_DISABLED))) + return -EINVAL; + + err =3D mnt_want_write_file(filp); + if (err) + return err; + + inode_lock(inode); + + if (!IS_DEVICE_ALIASING(inode)) { + err =3D -EINVAL; + goto out_inode_unlock; + } + + if (F2FS_HAS_BLOCKS(inode)) { + err =3D 0; + goto out_inode_unlock; + } + + for (i =3D 1; i < sbi->s_ndevs; i++) { + char *name =3D strrchr(FDEV(i).path, '/'); + + name =3D name ? name + 1 : FDEV(i).path; + if (!strcmp(name, filp->f_path.dentry->d_name.name)) { + ei.blk =3D FDEV(i).start_blk; + ei.len =3D FDEV(i).total_segments << sbi->log_blocks_per_seg; + ei.fofs =3D 0; + break; + } + } + + if (i =3D=3D sbi->s_ndevs) { + f2fs_warn(sbi, "device alias file (%s, ino=3D%llu) has no matching devic= e", + filp->f_path.dentry->d_name.name, inode->i_ino); + set_sbi_flag(sbi, SBI_NEED_FSCK); + f2fs_handle_error(sbi, ERROR_CORRUPTED_INODE); + err =3D -EFSCORRUPTED; + goto out_inode_unlock; + } + + f2fs_down_write_trace(&sbi->gc_lock, &glc); + f2fs_lock_op(sbi, &lc); + + if (unlikely(is_sbi_flag_set(sbi, SBI_CP_DISABLED))) { + err =3D -EINVAL; + } else { + count =3D ei.len; + err =3D inc_valid_block_count(sbi, inode, &count, false); + } + if (err) { + f2fs_unlock_op(sbi, &lc); + f2fs_up_write_trace(&sbi->gc_lock, &glc); + goto out_inode_unlock; + } + + spin_lock(&FREE_I(sbi)->segmap_lock); + FDEV(f2fs_target_device_index(sbi, ei.blk)).is_reserving =3D true; + spin_unlock(&FREE_I(sbi)->segmap_lock); + + start =3D GET_SEGNO(sbi, ei.blk); + end =3D GET_SEGNO(sbi, ei.blk + ei.len - 1); + + /* Reset the victim information to prevent GC from targeting the range */ + f2fs_reset_gc_victim_resource(sbi, start, end); + + /* Move out cursegs from the target range */ + for (type =3D CURSEG_HOT_DATA; type < NR_CURSEG_PERSIST_TYPE; type++) { + err =3D f2fs_allocate_segment_for_resize(sbi, type, start, end); + if (err) { + f2fs_unlock_op(sbi, &lc); + goto out_gc_unlock; + } + } + + f2fs_unlock_op(sbi, &lc); + f2fs_up_write_trace(&sbi->gc_lock, &glc); + + /* Write checkpoint synchronously to flush all pending writes and free sp= ace */ + err =3D f2fs_write_checkpoint(sbi, &cpc); + if (err) { + f2fs_down_write_trace(&sbi->gc_lock, &glc); + goto out_gc_unlock; + } + + /* Re-acquire gc_lock and cp_rwsem read lock for the entire range GC */ + f2fs_down_write_trace(&sbi->gc_lock, &glc); + f2fs_lock_op(sbi, &lc); + + if (unlikely(is_sbi_flag_set(sbi, SBI_CP_DISABLED))) { + err =3D -EINVAL; + f2fs_unlock_op(sbi, &lc); + goto out_gc_unlock; + } + + /* do GC to move out valid blocks in the range all at once! */ + err =3D f2fs_gc_range(sbi, start, end, false, 0); + if (err) { + f2fs_unlock_op(sbi, &lc); + goto out_gc_unlock; + } + + if (et) { + write_lock(&et->lock); + et->largest =3D ei; + write_unlock(&et->lock); + } + clear_inode_flag(inode, FI_NO_EXTENT); + + f2fs_reserve_device_alias(sbi, ei.blk, ei.len); + + i_size_write(inode, (loff_t)ei.len << PAGE_SHIFT); + f2fs_update_inode_page(inode); + + spin_lock(&FREE_I(sbi)->segmap_lock); + FDEV(f2fs_target_device_index(sbi, ei.blk)).is_reserving =3D false; + spin_unlock(&FREE_I(sbi)->segmap_lock); + + f2fs_unlock_op(sbi, &lc); + f2fs_up_write_trace(&sbi->gc_lock, &glc); + + inode_unlock(inode); + mnt_drop_write_file(filp); + + err =3D f2fs_write_checkpoint(sbi, &cpc); + return err; + +out_gc_unlock: + spin_lock(&FREE_I(sbi)->segmap_lock); + FDEV(f2fs_target_device_index(sbi, ei.blk)).is_reserving =3D false; + spin_unlock(&FREE_I(sbi)->segmap_lock); + f2fs_up_write_trace(&sbi->gc_lock, &glc); + + /* + * Put successfully GC'ed segments back into PRE list so checkpoint + * commits and frees them! + */ + f2fs_lock_op(sbi, &lc); + for (segno =3D start; segno <=3D end; segno++) { + if (get_valid_blocks(sbi, segno, false) =3D=3D 0) { + mutex_lock(&DIRTY_I(sbi)->seglist_lock); + if (!test_and_set_bit(segno, DIRTY_I(sbi)->dirty_segmap[PRE])) + DIRTY_I(sbi)->nr_dirty[PRE]++; + mutex_unlock(&DIRTY_I(sbi)->seglist_lock); + } + } + count =3D ei.len; + dec_valid_block_count(sbi, inode, count); + f2fs_unlock_op(sbi, &lc); + + inode_unlock(inode); + mnt_drop_write_file(filp); + + f2fs_write_checkpoint(sbi, &cpc); + return err; + +out_inode_unlock: + inode_unlock(inode); + mnt_drop_write_file(filp); + return err; +} + +static int f2fs_ioc_release_dev_alias(struct file *filp) +{ + struct inode *inode =3D file_inode(filp); + struct f2fs_sb_info *sbi =3D F2FS_I_SB(inode); + struct extent_tree *et =3D F2FS_I(inode)->extent_tree[EX_READ]; + struct extent_info ei =3D {0, }; + struct cp_control cpc =3D { CP_SYNC, 0, 0, 0 }; + struct f2fs_lock_context lc, glc; + int err; + + if (!capable(CAP_SYS_ADMIN)) + return -EPERM; + + if (unlikely(is_sbi_flag_set(sbi, SBI_CP_DISABLED))) + return -EINVAL; + + err =3D mnt_want_write_file(filp); + if (err) + return err; + + inode_lock(inode); + + if (!IS_DEVICE_ALIASING(inode)) { + err =3D -EINVAL; + goto out_inode_unlock; + } + + if (!F2FS_HAS_BLOCKS(inode)) { + err =3D 0; + goto out_inode_unlock; + } + + err =3D filemap_write_and_wait(inode->i_mapping); + if (err) + goto out_inode_unlock; + + read_lock(&et->lock); + ei =3D et->largest; + read_unlock(&et->lock); + + f2fs_down_write_trace(&sbi->gc_lock, &glc); + f2fs_lock_op(sbi, &lc); + + if (unlikely(is_sbi_flag_set(sbi, SBI_CP_DISABLED))) { + err =3D -EINVAL; + f2fs_unlock_op(sbi, &lc); + f2fs_up_write_trace(&sbi->gc_lock, &glc); + goto out_inode_unlock; + } + + truncate_setsize(inode, 0); + + err =3D f2fs_truncate_blocks(inode, 0, false); + if (err) { + i_size_write(inode, (loff_t)ei.len << PAGE_SHIFT); + f2fs_unlock_op(sbi, &lc); + f2fs_up_write_trace(&sbi->gc_lock, &glc); + goto out_inode_unlock; + } + + f2fs_update_inode_page(inode); + + f2fs_unlock_op(sbi, &lc); + f2fs_up_write_trace(&sbi->gc_lock, &glc); + + inode_unlock(inode); + mnt_drop_write_file(filp); + + err =3D f2fs_write_checkpoint(sbi, &cpc); + return err; + +out_inode_unlock: + inode_unlock(inode); + mnt_drop_write_file(filp); + return err; +} + static int f2fs_ioc_io_prio(struct file *filp, unsigned long arg) { struct inode *inode =3D file_inode(filp); @@ -4742,8 +5016,14 @@ static long __f2fs_ioctl(struct file *filp, unsigned= int cmd, unsigned long arg) return f2fs_ioc_compress_file(filp); case F2FS_IOC_GET_DEV_ALIAS_FILE: return f2fs_ioc_get_dev_alias_file(filp, arg); + case F2FS_IOC_GET_DEV_ALIAS_STATUS: + return f2fs_ioc_get_dev_alias_status(filp, arg); case F2FS_IOC_IO_PRIO: return f2fs_ioc_io_prio(filp, arg); + case F2FS_IOC_RESERVE_DEV_ALIAS: + return f2fs_ioc_reserve_dev_alias(filp); + case F2FS_IOC_RELEASE_DEV_ALIAS: + return f2fs_ioc_release_dev_alias(filp); default: return -ENOTTY; } @@ -5530,7 +5810,10 @@ long f2fs_compat_ioctl(struct file *file, unsigned i= nt cmd, unsigned long arg) case F2FS_IOC_DECOMPRESS_FILE: case F2FS_IOC_COMPRESS_FILE: case F2FS_IOC_GET_DEV_ALIAS_FILE: + case F2FS_IOC_GET_DEV_ALIAS_STATUS: case F2FS_IOC_IO_PRIO: + case F2FS_IOC_RESERVE_DEV_ALIAS: + case F2FS_IOC_RELEASE_DEV_ALIAS: break; default: return -ENOIOCTLCMD; diff --git a/fs/f2fs/gc.c b/fs/f2fs/gc.c index ffaa7ba76a1b..93bcb35a5b5d 100644 --- a/fs/f2fs/gc.c +++ b/fs/f2fs/gc.c @@ -2197,29 +2197,37 @@ int f2fs_gc_range(struct f2fs_sb_info *sbi, return 0; } =20 +void f2fs_reset_gc_victim_resource(struct f2fs_sb_info *sbi, + unsigned int start, unsigned int end) +{ + int i; + + mutex_lock(&DIRTY_I(sbi)->seglist_lock); + for (i =3D 0; i < MAX_GC_POLICY; i++) + if (SIT_I(sbi)->last_victim[i] >=3D start && + SIT_I(sbi)->last_victim[i] <=3D end) + SIT_I(sbi)->last_victim[i] =3D 0; + + for (i =3D BG_GC; i <=3D FG_GC; i++) + if (sbi->next_victim_seg[i] >=3D start && + sbi->next_victim_seg[i] <=3D end) + sbi->next_victim_seg[i] =3D NULL_SEGNO; + mutex_unlock(&DIRTY_I(sbi)->seglist_lock); +} + static int free_segment_range(struct f2fs_sb_info *sbi, unsigned int secs, bool dry_run) { unsigned int next_inuse, start, end; struct cp_control cpc =3D { CP_RESIZE, 0, 0, 0 }; - int gc_mode, gc_type; int err =3D 0; int type; =20 - /* Force block allocation for GC */ MAIN_SECS(sbi) -=3D secs; start =3D MAIN_SECS(sbi) * SEGS_PER_SEC(sbi); end =3D MAIN_SEGS(sbi) - 1; =20 - mutex_lock(&DIRTY_I(sbi)->seglist_lock); - for (gc_mode =3D 0; gc_mode < MAX_GC_POLICY; gc_mode++) - if (SIT_I(sbi)->last_victim[gc_mode] >=3D start) - SIT_I(sbi)->last_victim[gc_mode] =3D 0; - - for (gc_type =3D BG_GC; gc_type <=3D FG_GC; gc_type++) - if (sbi->next_victim_seg[gc_type] >=3D start) - sbi->next_victim_seg[gc_type] =3D NULL_SEGNO; - mutex_unlock(&DIRTY_I(sbi)->seglist_lock); + f2fs_reset_gc_victim_resource(sbi, start, end); =20 /* Move out cursegs from the target range */ for (type =3D CURSEG_HOT_DATA; type < NR_CURSEG_PERSIST_TYPE; type++) { diff --git a/fs/f2fs/namei.c b/fs/f2fs/namei.c index cac03b8e91a1..8c3b57987f6c 100644 --- a/fs/f2fs/namei.c +++ b/fs/f2fs/namei.c @@ -425,6 +425,9 @@ static int f2fs_link(struct dentry *old_dentry, struct = inode *dir, if (!f2fs_is_checkpoint_ready(sbi)) return -ENOSPC; =20 + if (IS_DEVICE_ALIASING(inode)) + return -EPERM; + err =3D fscrypt_prepare_link(old_dentry, dir, dentry); if (err) return err; @@ -568,6 +571,9 @@ static int f2fs_unlink(struct inode *dir, struct dentry= *dentry) =20 trace_f2fs_unlink_enter(dir, dentry); =20 + if (IS_DEVICE_ALIASING(inode)) + return -EPERM; + if (unlikely(f2fs_cp_error(sbi))) { err =3D -EIO; goto out; @@ -946,6 +952,9 @@ static int f2fs_rename(struct mnt_idmap *idmap, struct = inode *old_dir, bool old_is_dir =3D S_ISDIR(old_inode->i_mode); int err; =20 + if (IS_DEVICE_ALIASING(old_inode)) + return -EPERM; + if (unlikely(f2fs_cp_error(sbi))) return -EIO; if (!f2fs_is_checkpoint_ready(sbi)) @@ -1016,6 +1025,8 @@ static int f2fs_rename(struct mnt_idmap *idmap, struc= t inode *old_dir, } =20 if (new_inode) { + if (IS_DEVICE_ALIASING(new_inode)) + return -EPERM; =20 err =3D -ENOTEMPTY; if (old_is_dir && !f2fs_empty_dir(new_inode)) @@ -1143,6 +1154,9 @@ static int f2fs_cross_rename(struct inode *old_dir, s= truct dentry *old_dentry, int old_nlink =3D 0, new_nlink =3D 0; int err; =20 + if (IS_DEVICE_ALIASING(old_inode) || IS_DEVICE_ALIASING(new_inode)) + return -EPERM; + if (unlikely(f2fs_cp_error(sbi))) return -EIO; if (!f2fs_is_checkpoint_ready(sbi)) diff --git a/fs/f2fs/segment.c b/fs/f2fs/segment.c index d71ddb3ee918..1a986f1884f2 100644 --- a/fs/f2fs/segment.c +++ b/fs/f2fs/segment.c @@ -2502,35 +2502,42 @@ static int update_sit_entry_for_alloc(struct f2fs_s= b_info *sbi, struct seg_entry unsigned int segno, block_t blkaddr, unsigned int offset, int del) { bool exist; + int del_count =3D del; + int i; =20 - exist =3D f2fs_test_and_set_bit(offset, se->cur_valid_map); - if (unlikely(exist)) { - f2fs_err(sbi, "Bitmap was wrongly set, blk:%u", blkaddr); - f2fs_bug_on(sbi, 1); - se->valid_blocks--; - del =3D 0; - } + f2fs_bug_on(sbi, GET_SEGNO(sbi, blkaddr) !=3D GET_SEGNO(sbi, blkaddr + de= l_count - 1)); =20 - if (f2fs_block_unit_discard(sbi) && - !f2fs_test_and_set_bit(offset, se->discard_map)) - sbi->discard_blks--; + for (i =3D 0; i < del_count; i++) { + exist =3D f2fs_test_and_set_bit(offset + i, se->cur_valid_map); + if (unlikely(exist)) { + f2fs_err(sbi, "Bitmap was wrongly set, blk:%u", blkaddr + i); + f2fs_bug_on(sbi, 1); + se->valid_blocks--; + del -=3D 1; + continue; + } =20 - /* - * SSR should never reuse block which is checkpointed - * or newly invalidated. - */ - if (!is_sbi_flag_set(sbi, SBI_CP_DISABLED)) { - if (!f2fs_test_and_set_bit(offset, se->ckpt_valid_map)) { - se->ckpt_valid_blocks++; - if (__is_large_section(sbi)) - get_sec_entry(sbi, segno)->ckpt_valid_blocks++; + if (f2fs_block_unit_discard(sbi) && + !f2fs_test_and_set_bit(offset + i, se->discard_map)) + sbi->discard_blks--; + + /* + * SSR should never reuse block which is checkpointed + * or newly invalidated. + */ + if (!is_sbi_flag_set(sbi, SBI_CP_DISABLED)) { + if (!f2fs_test_and_set_bit(offset + i, se->ckpt_valid_map)) { + se->ckpt_valid_blocks++; + if (__is_large_section(sbi)) + get_sec_entry(sbi, segno)->ckpt_valid_blocks++; + } } - } =20 - if (!f2fs_test_bit(offset, se->ckpt_valid_map)) { - se->ckpt_valid_blocks +=3D del; - if (__is_large_section(sbi)) - get_sec_entry(sbi, segno)->ckpt_valid_blocks +=3D del; + if (!f2fs_test_bit(offset + i, se->ckpt_valid_map)) { + se->ckpt_valid_blocks +=3D 1; + if (__is_large_section(sbi)) + get_sec_entry(sbi, segno)->ckpt_valid_blocks +=3D 1; + } } =20 if (__is_large_section(sbi)) @@ -2585,9 +2592,14 @@ void f2fs_invalidate_blocks(struct f2fs_sb_info *sbi= , block_t addr, unsigned int segno =3D GET_SEGNO(sbi, addr); struct sit_info *sit_i =3D SIT_I(sbi); block_t addr_start =3D addr, addr_end =3D addr + len - 1; - unsigned int seg_num =3D GET_SEGNO(sbi, addr_end) - segno + 1; + unsigned int seg_num; unsigned int i =3D 1, max_blocks =3D sbi->blocks_per_seg, cnt; =20 + if (len =3D=3D 0) + return; + + seg_num =3D GET_SEGNO(sbi, addr_end) - segno + 1; + f2fs_bug_on(sbi, addr =3D=3D NULL_ADDR); if (addr =3D=3D NEW_ADDR || addr =3D=3D COMPRESS_ADDR) return; @@ -2620,6 +2632,52 @@ void f2fs_invalidate_blocks(struct f2fs_sb_info *sbi= , block_t addr, up_write(&sit_i->sentry_lock); } =20 +void f2fs_reserve_device_alias(struct f2fs_sb_info *sbi, block_t addr, + unsigned int len) +{ + unsigned int segno =3D GET_SEGNO(sbi, addr); + struct sit_info *sit_i =3D SIT_I(sbi); + block_t addr_start =3D addr, addr_end =3D addr + len - 1; + unsigned int seg_num; + unsigned int i =3D 1, max_blocks =3D sbi->blocks_per_seg, cnt; + + if (len =3D=3D 0) + return; + + seg_num =3D GET_SEGNO(sbi, addr_end) - segno + 1; + + down_write(&sit_i->sentry_lock); + + if (seg_num =3D=3D 1) + cnt =3D len; + else + cnt =3D max_blocks - GET_BLKOFF_FROM_SEG0(sbi, addr); + + do { + update_segment_mtime(sbi, addr_start, 0); + update_sit_entry(sbi, addr_start, cnt); + __set_test_and_inuse(sbi, segno); + + /* Remove the segment from PRE (prefree) to prevent checkpoint from free= ing it! */ + mutex_lock(&DIRTY_I(sbi)->seglist_lock); + if (test_and_clear_bit(segno, DIRTY_I(sbi)->dirty_segmap[PRE])) + DIRTY_I(sbi)->nr_dirty[PRE]--; + mutex_unlock(&DIRTY_I(sbi)->seglist_lock); + + /* add it into dirty seglist */ + locate_dirty_segment(sbi, segno); + + /* update @addr_start and @cnt and @segno */ + addr_start =3D START_BLOCK(sbi, ++segno); + if (++i =3D=3D seg_num) + cnt =3D GET_BLKOFF_FROM_SEG0(sbi, addr_end) + 1; + else + cnt =3D max_blocks; + } while (i <=3D seg_num); + + up_write(&sit_i->sentry_lock); +} + bool f2fs_is_checkpointed_data(struct f2fs_sb_info *sbi, block_t blkaddr) { struct sit_info *sit_i =3D SIT_I(sbi); @@ -2758,8 +2816,13 @@ static int is_next_segment_free(struct f2fs_sb_info = *sbi, unsigned int segno =3D curseg->segno + 1; struct free_segmap_info *free_i =3D FREE_I(sbi); =20 - if (segno < MAIN_SEGS(sbi) && segno % SEGS_PER_SEC(sbi)) + if (segno < MAIN_SEGS(sbi) && segno % SEGS_PER_SEC(sbi)) { + int devi =3D f2fs_target_device_index(sbi, START_BLOCK(sbi, segno)); + + if (f2fs_dev_is_reserving(sbi, devi)) + return 0; return !test_bit(segno, free_i->free_segmap); + } return 0; } =20 @@ -2778,6 +2841,7 @@ static int get_new_segment(struct f2fs_sb_info *sbi, unsigned int alloc_policy =3D sbi->allocate_section_policy; unsigned int alloc_hint =3D sbi->allocate_section_hint; bool init =3D true; + bool looped =3D false; int i; int ret =3D 0; =20 @@ -2791,8 +2855,13 @@ static int get_new_segment(struct f2fs_sb_info *sbi, if (!new_sec && ((*newseg + 1) % SEGS_PER_SEC(sbi))) { segno =3D find_next_zero_bit(free_i->free_segmap, GET_SEG_FROM_SEC(sbi, hint + 1), *newseg + 1); - if (segno < GET_SEG_FROM_SEC(sbi, hint + 1)) + if (segno < GET_SEG_FROM_SEC(sbi, hint + 1)) { + int devi =3D f2fs_target_device_index(sbi, START_BLOCK(sbi, segno)); + + if (f2fs_dev_is_alloc_blocked(sbi, devi, pinning)) + goto find_other_zone; goto got_it; + } } =20 #ifdef CONFIG_BLK_DEV_ZONED @@ -2828,33 +2897,50 @@ static int get_new_segment(struct f2fs_sb_info *sbi, find_other_zone: secno =3D find_next_zero_bit(free_i->free_secmap, MAIN_SECS(sbi), hint); =20 -#ifdef CONFIG_BLK_DEV_ZONED - if (secno >=3D MAIN_SECS(sbi) && f2fs_sb_has_blkzoned(sbi)) { - /* Write only to sequential zones */ - if (sbi->blkzone_alloc_policy =3D=3D BLKZONE_ALLOC_ONLY_SEQ) { - hint =3D GET_SEC_FROM_SEG(sbi, sbi->first_seq_zone_segno); - secno =3D find_next_zero_bit(free_i->free_secmap, MAIN_SECS(sbi), hint); - } else - secno =3D find_first_zero_bit(free_i->free_secmap, - MAIN_SECS(sbi)); - if (secno >=3D MAIN_SECS(sbi)) { - ret =3D -ENOSPC; - f2fs_bug_on(sbi, 1); - goto out_unlock; - } - } -#endif - if (secno >=3D MAIN_SECS(sbi)) { - secno =3D find_first_zero_bit(free_i->free_secmap, - MAIN_SECS(sbi)); - if (secno >=3D MAIN_SECS(sbi)) { + if (looped) { ret =3D -ENOSPC; f2fs_bug_on(sbi, !pinning); goto out_unlock; } +#ifdef CONFIG_BLK_DEV_ZONED + /* Write only to sequential zones */ + if (f2fs_sb_has_blkzoned(sbi) && + sbi->blkzone_alloc_policy =3D=3D BLKZONE_ALLOC_ONLY_SEQ) + hint =3D GET_SEC_FROM_SEG(sbi, sbi->first_seq_zone_segno); + else +#endif + hint =3D 0; + looped =3D true; + goto find_other_zone; } + segno =3D GET_SEG_FROM_SEC(sbi, secno); + + if (f2fs_sb_has_device_alias(sbi) && f2fs_is_multi_device(sbi)) { + int devi =3D f2fs_target_device_index(sbi, START_BLOCK(sbi, segno)); + + if (f2fs_dev_is_alloc_blocked(sbi, devi, pinning)) { + unsigned int end_segno; + + while (devi < sbi->s_ndevs && + f2fs_dev_is_alloc_blocked(sbi, devi, pinning)) { + block_t next_blk; + + end_segno =3D GET_SEGNO(sbi, FDEV(devi).end_blk); + hint =3D GET_SEC_FROM_SEG(sbi, end_segno) + 1; + + if (hint >=3D MAIN_SECS(sbi) || ++devi >=3D sbi->s_ndevs) + break; + + next_blk =3D START_BLOCK(sbi, GET_SEG_FROM_SEC(sbi, hint)); + if (next_blk < FDEV(devi).start_blk || + next_blk > FDEV(devi).end_blk) + break; + } + goto find_other_zone; + } + } zoneno =3D GET_ZONE_FROM_SEC(sbi, secno); =20 /* give up on finding another zone */ diff --git a/fs/f2fs/segment.h b/fs/f2fs/segment.h index b0c06b3580b4..259e4a29b494 100644 --- a/fs/f2fs/segment.h +++ b/fs/f2fs/segment.h @@ -954,10 +954,36 @@ static inline block_t sum_blk_addr(struct f2fs_sb_inf= o *sbi, int base, int type) - (base + 1) + type; } =20 +static inline bool f2fs_dev_is_reserving(struct f2fs_sb_info *sbi, int dev= i) +{ + if (!f2fs_sb_has_device_alias(sbi) || !f2fs_is_multi_device(sbi)) + return false; + return FDEV(devi).is_reserving; +} + +static inline bool f2fs_dev_is_alloc_blocked(struct f2fs_sb_info *sbi, + int devi, bool pinning) +{ + if (!f2fs_sb_has_device_alias(sbi) || !f2fs_is_multi_device(sbi)) + return false; + return (pinning && FDEV(devi).has_alias) || FDEV(devi).is_reserving; +} + static inline bool sec_usage_check(struct f2fs_sb_info *sbi, unsigned int = secno) { if (is_cursec(sbi, secno) || (sbi->cur_victim_sec =3D=3D secno)) return true; + if (f2fs_sb_has_device_alias(sbi) && f2fs_is_multi_device(sbi)) { + int i; + block_t start_blk =3D START_BLOCK(sbi, GET_SEG_FROM_SEC(sbi, secno)); + + for (i =3D 0; i < sbi->s_ndevs; i++) { + if (f2fs_dev_is_reserving(sbi, i) && + start_blk >=3D FDEV(i).start_blk && + start_blk <=3D FDEV(i).end_blk) + return true; + } + } return false; } =20 diff --git a/fs/f2fs/super.c b/fs/f2fs/super.c index c448d992ff2a..8359eed903be 100644 --- a/fs/f2fs/super.c +++ b/fs/f2fs/super.c @@ -5001,6 +5001,38 @@ static void f2fs_tuning_parameters(struct f2fs_sb_in= fo *sbi) sbi->readdir_ra =3D true; } =20 +static void f2fs_restore_device_alias(struct f2fs_sb_info *sbi) +{ + struct inode *root =3D d_inode(sbi->sb->s_root); + struct f2fs_dir_entry *de; + struct folio *folio; + int i; + + if (!f2fs_sb_has_device_alias(sbi)) + return; + + for (i =3D 1; i < sbi->s_ndevs; i++) { + char *name =3D strrchr(FDEV(i).path, '/'); + struct qstr qstr; + + name =3D name ? name + 1 : FDEV(i).path; + qstr.name =3D name; + qstr.len =3D strlen(name); + + de =3D f2fs_find_entry(root, &qstr, &folio); + if (de) { + struct inode *inode =3D f2fs_iget(sbi->sb, le32_to_cpu(de->ino)); + + if (!IS_ERR(inode)) { + if (IS_DEVICE_ALIASING(inode)) + FDEV(i).has_alias =3D true; + iput(inode); + } + f2fs_folio_put(folio, 0); + } + } +} + static int f2fs_fill_super(struct super_block *sb, struct fs_context *fc) { struct f2fs_fs_context *ctx =3D fc->fs_private; @@ -5436,6 +5468,8 @@ static int f2fs_fill_super(struct super_block *sb, st= ruct fs_context *fc) f2fs_update_time(sbi, REQ_TIME); clear_sbi_flag(sbi, SBI_CP_DISABLED_QUICK); =20 + f2fs_restore_device_alias(sbi); + sbi->umount_lock_holder =3D NULL; return 0; =20 diff --git a/include/uapi/linux/f2fs.h b/include/uapi/linux/f2fs.h index 795e26258355..4409ada2fecb 100644 --- a/include/uapi/linux/f2fs.h +++ b/include/uapi/linux/f2fs.h @@ -45,6 +45,9 @@ #define F2FS_IOC_START_ATOMIC_REPLACE _IO(F2FS_IOCTL_MAGIC, 25) #define F2FS_IOC_GET_DEV_ALIAS_FILE _IOR(F2FS_IOCTL_MAGIC, 26, __u32) #define F2FS_IOC_IO_PRIO _IOW(F2FS_IOCTL_MAGIC, 27, __u32) +#define F2FS_IOC_RESERVE_DEV_ALIAS _IO(F2FS_IOCTL_MAGIC, 28) +#define F2FS_IOC_RELEASE_DEV_ALIAS _IO(F2FS_IOCTL_MAGIC, 29) +#define F2FS_IOC_GET_DEV_ALIAS_STATUS _IOR(F2FS_IOCTL_MAGIC, 30, __u32) =20 /* * should be same as XFS_IOC_GOINGDOWN. @@ -70,6 +73,10 @@ enum { F2FS_IOPRIO_MAX, }; =20 +/* for F2FS_IOC_GET_DEV_ALIAS_STATUS */ +#define F2FS_DEV_ALIAS_STATUS_RELEASED 0 +#define F2FS_DEV_ALIAS_STATUS_RESERVED 1 + struct f2fs_gc_range { __u32 sync; __u64 start; --=20 2.55.0.795.g602f6c329a-goog