From nobody Fri Jul 24 21:52:37 2026 Received: from mail-pl1-f169.google.com (mail-pl1-f169.google.com [209.85.214.169]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id A235D3D967F for ; Thu, 23 Jul 2026 20:20:16 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.214.169 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784838021; cv=none; b=rdEJSf8F33MqEMW6SgdZi7dGoJuwsPkrI/NYSDH77Wh1u8+/p9k7GXB7wNcXjXuPG5Xd+lID/7gr/NI/iXEB3Esz3cR9YNR844Dt/+blXr52VxNmtcvDdFu2bkmVhJ/xEILBLjSigXO/Wgzim0RHUL3dV8+1jc2X8+sygv2vGqI= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784838021; c=relaxed/simple; bh=fMweb792GArpTQRxFXPvB18+644pJWlnQRRVaNMVzOU=; h=From:To:Cc:Subject:Date:Message-ID:MIME-Version; b=PjPeYCIgRWANEy9DcWUUvSQEnhcUafd3Bhvm8jfFORordClP+WuY45En17OhkNK1+HSUGQaBjR4zbWwNdvpmj9j6KNZyFY8pt+gcH9C+7j0uAON6Rmg1x83XnXDCQZT+BBnAZZNll+u8sgUKfM9Xl8CdzfDsaYKDSv5kVVfgevs= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=bfo3zbDv; arc=none smtp.client-ip=209.85.214.169 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="bfo3zbDv" Received: by mail-pl1-f169.google.com with SMTP id d9443c01a7336-2caea3f742bso14292895ad.0 for ; Thu, 23 Jul 2026 13:20:15 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1784838015; x=1785442815; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:message-id:date:subject:cc :to:from:from:to:cc:subject:date:message-id:reply-to:content-type; bh=OBLMBIlNjKE3HOa02vgj2Sd8aW1XFgIO2esIkhLqebI=; b=bfo3zbDvKlaqkZg2sv8I3o8nMHipeJd3zNIlC/dMdnXiocWhyja4T/DOoOM9hTukWq TsGTiIHlzhIfiMmDNVG2hd4GWd8r20SqTRJbD29733X/N1FGIepejzyFJ4TQ4jDLzS3u lvI3AfZU420T7rbPmwCglheOgwOBvOWL0n2uiYAmK1zOgJUY4AwMblhQwPXl9AwUvsZ4 PTwrbVp65nHHKyBK/iQmMrRLXLkEB7Cdh2x+JYcAst3U9eFSi27DTv0wOER1cXPsMyZi VzqvSEPk7ZacPVQgOpBgcS26yajyTe7n+YeVXGwVBr3/s7kyAsg0m2UH8aRcad5I/reJ 1HNQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1784838015; x=1785442815; h=content-transfer-encoding:mime-version:message-id:date:subject:cc :to:from:x-gm-gg:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=OBLMBIlNjKE3HOa02vgj2Sd8aW1XFgIO2esIkhLqebI=; b=p/0unaUr/e9u/i72WOfrW5tft9IC28DRxB4aRKRk8iwwMRLMkrEtwwozHMe6H0LyV6 m67MeMnAKSlTtmqFHMrxfhmQ9UyKmeqE39RKN8ByHUiiqCsOLd23kR8b+HOT2iTY4gYF Do4tk7aCr9PFHf/RzS59wrhokbT0zwWZGa5KvQQMGugzRZfusHknSUciaNK0A9pQEYey wCAbE1Pmw9P0/s1v0B3fHh91qmN0aTgpjI/1ZHBTAtUIKvnMLl/YDiktQwI6/WH3rQlg x5OwB9JYH5bpoFJlEZAdeK6CIf6Lw+7monbUAmeJjg2sGfasZrCfAJiZoTMR7x7aYqyk Z7zg== X-Gm-Message-State: AOJu0YzmLJQxnVxp9rPT12cShok0vZZ7inR1h2Q3t4M5aVTZ2pDkWr9o W5ao3Lxu3zf/WMYu14zyG7mgcoD7NTu8vcIk1RqUlL0IMyLCURBaofYOTvdu6A== X-Gm-Gg: AR+sD13WMLODXmIhtAiByFBFtyTWx1jh+pQAOWYQnJXNJt7xtsRk7jIJvTYUeJ+U3db Ywhtt88To0Dc7CHodeRn0rdL28v/Ov4JdcWtgBnWGqYa6rkXMZmarkx7lUaAQJ8baqR5cYuaZMp ozR74bmPTSLEuNl3NEyV2XT9xOSP7kwq66w6o8yjaTwp27sXZOq0b1bchk9okVI7SN0fkr8qrdk dizLFyARig0xeD4FQrL97NOEXyUIQl/D5CrJPrLgyXAy64cl5yL5FQAqw5aqTBzzhPq6EWme64H HKjvP27j2n7zV2pzaqCAmGUmypnMybrrtN60YFEXwjpeld4oQ7ILxsz4YDSiwAIDZ4D0qoxOZG+ Z63gOTOx3LSHk6iAx0kVH11V0aRvo8LmIQ2QFHlnSuwctCmLw2hmGC9ZTQP+58WGnxYMr/iTbgk Q+LWDxQfdUdC5L7Jey3M4CdFK8kRrK+ZvBRDekS4XMU49OjWHiAUznf6YP3tIxFO2KMvec+M/E2 1UMMxibHCfqGeZXqiz7QTQlBm8hSbNhLiM+T+co X-Received: by 2002:a05:6a21:648c:b0:3c0:9c19:6586 with SMTP id adf61e73a8af0-3c44b25f194mr5350815637.64.1784838014481; Thu, 23 Jul 2026 13:20:14 -0700 (PDT) Received: from daehojeong-desktop.mtv.corp.google.com ([2a00:79e0:2e7c:8:8f80:8150:b723:2d96]) by smtp.gmail.com with ESMTPSA id 5a478bee46e88-3147dc6781dsm24591070eec.9.2026.07.23.13.20.13 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Thu, 23 Jul 2026 13:20:13 -0700 (PDT) From: Daeho Jeong To: linux-kernel@vger.kernel.org, linux-f2fs-devel@lists.sourceforge.net, kernel-team@android.com Cc: Daeho Jeong Subject: [PATCH v5] f2fs: support dynamic reserve/release for device aliasing Date: Thu, 23 Jul 2026 13:20:08 -0700 Message-ID: <20260723202008.2581659-1-daeho43@gmail.com> X-Mailer: git-send-email 2.55.0.229.g6434b31f56-goog Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Daeho Jeong This patch adds a dynamic management feature to the existing device aliasing functionality. It allows users to dynamically reserve or release specific devices from the filesystem's free pool at runtime through new ioctls. To support this, three new ioctls are introduced: - F2FS_IOC_RESERVE_DEV_ALIAS: This reclaims the space occupied by a device aliasing file. It first performs a capacity check, resets GC victim information for the target range, marks the segments as in-use to prevent new allocations, and then triggers GC to migrate existing valid data out of the range. Finally, it reserves these blocks in the SIT to effectively exclude the device from the usable capacity. - F2FS_IOC_RELEASE_DEV_ALIAS: This releases the reserved space of a previously reserved device aliasing file. It truncates the blocks associated with the file, which makes them available for general filesystem allocation again. - F2FS_IOC_GET_DEV_ALIAS_STATUS: This retrieves the current aliasing status of a device aliasing file, returning whether the file is released (inactive alias) or reserved (active alias, with blocks fully allocated on the device). Signed-off-by: Daeho Jeong --- v5: prevented new segment allocation on devices undergoing reservation or with active aliases. used in-memory struct to reserve blocks for device aliasing. cleared prefree (PRE) dirty segment bitmap. v4: renamed interfaces. fixed race conditions between checkpoint=3Ddisable mount and ioctls. refactored segment reservation part. modified lock usage. v3: add CAP_SYS_ADMIN and checkpoint=3Ddisabled check. remove a f2fs specific flag exposed with getflags. v2: prevent operations during checkpoint=3Ddisabled. --- --- Documentation/filesystems/f2fs.rst | 35 ++++ fs/f2fs/f2fs.h | 10 +- fs/f2fs/file.c | 278 ++++++++++++++++++++++++++++- fs/f2fs/gc.c | 30 ++-- fs/f2fs/namei.c | 14 ++ fs/f2fs/segment.c | 174 +++++++++++++----- fs/f2fs/segment.h | 22 +++ fs/f2fs/super.c | 35 ++++ include/uapi/linux/f2fs.h | 7 + 9 files changed, 542 insertions(+), 63 deletions(-) diff --git a/Documentation/filesystems/f2fs.rst b/Documentation/filesystems= /f2fs.rst index 8c4a14ae444f..1a5fd4afe609 100644 --- a/Documentation/filesystems/f2fs.rst +++ b/Documentation/filesystems/f2fs.rst @@ -1045,6 +1045,41 @@ So, the key idea is, user can do any file operations= on /dev/vdc, and reclaim the space after the use, while the space is counted as /data. That doesn't require modifying partition size and filesystem format. =20 +Dynamic Device Aliasing Management +---------------------------------- + +In addition to static device aliasing by deleting the aliasing file, F2FS +supports dynamic management of device aliasing. This mechanism allows the = system +to dynamically transition partition ownership between F2FS userdata and ex= ternal +entities (e.g., zRAM, raw partition) based on system requirements without +deleting the master aliasing file or requiring unmount/remount. + +The master aliasing file is created during the initial format of the file = system +and remains as a persistent control entity (ioctl gateway) in the root dir= ectory. + +- Partition Reservation (In-service to Aliased) + When a specific partition needs to be dedicated to external services (e.= g., zRAM), + a user can reserve the device alias range via ioctl. The kernel resets G= C victim + information for the target range, marks segments as in-use to prevent new + allocations, and triggers forced GC to migrate existing valid data out o= f the + range. Finally, it reserves these blocks in the SIT to effectively exclu= de the + device from the usable capacity. + +- Partition Release (Aliased to In-service) + When external usage concludes, the space is reclaimed not by deleting th= e file, + but through the release ioctl. The kernel truncates blocks associated wi= th + the file, releasing them back to general filesystem allocation. + +.. code-block:: + + # f2fs_io dev_alias release /mnt/f2fs/vdc.file + # df -h + /dev/vdb 64G 753M 64G 2% /mnt/f2fs + + # f2fs_io dev_alias reserve /mnt/f2fs/vdc.file + # df -h + /dev/vdb 64G 33G 32G 52% /mnt/f2fs + Per-file Read-Only Large Folio Support -------------------------------------- =20 diff --git a/fs/f2fs/f2fs.h b/fs/f2fs/f2fs.h index f1774d4e18d2..74c2d58ce271 100644 --- a/fs/f2fs/f2fs.h +++ b/fs/f2fs/f2fs.h @@ -1404,6 +1404,8 @@ struct f2fs_dev_info { unsigned int total_segments; block_t start_blk; block_t end_blk; + bool has_alias; + bool is_reserving; #ifdef CONFIG_BLK_DEV_ZONED unsigned int nr_blkz; /* Total number of zones */ unsigned long *blkz_seq; /* Bitmap indicating sequential zones */ @@ -1857,6 +1859,7 @@ struct f2fs_sb_info { block_t last_valid_block_count; /* for recovery */ block_t reserved_blocks; /* configurable reserved blocks */ block_t current_reserved_blocks; /* current reserved blocks */ + block_t alias_reserved_blocks; /* reserved blocks for device alias */ =20 /* Additional tracking for no checkpoint mode */ block_t unusable_block_count; /* # of blocks saved by last cp */ @@ -2559,7 +2562,8 @@ static inline unsigned int get_available_block_count(= struct f2fs_sb_info *sbi, block_t avail_user_block_count; =20 avail_user_block_count =3D sbi->user_block_count - - sbi->current_reserved_blocks; + sbi->current_reserved_blocks - + sbi->alias_reserved_blocks; =20 if (test_opt(sbi, RESERVE_ROOT) && !__allow_reserved_root(sbi, inode, cap= )) avail_user_block_count -=3D F2FS_OPTION(sbi).root_reserved_blocks; @@ -4010,6 +4014,8 @@ int f2fs_flush_device_cache(struct f2fs_sb_info *sbi); void f2fs_destroy_flush_cmd_control(struct f2fs_sb_info *sbi, bool free); void f2fs_invalidate_blocks(struct f2fs_sb_info *sbi, block_t addr, unsigned int len); +void f2fs_reserve_device_alias(struct f2fs_sb_info *sbi, block_t addr, + unsigned int len); bool f2fs_is_checkpointed_data(struct f2fs_sb_info *sbi, block_t blkaddr); int f2fs_start_discard_thread(struct f2fs_sb_info *sbi); void f2fs_drop_discard_cmd(struct f2fs_sb_info *sbi); @@ -4231,6 +4237,8 @@ void f2fs_build_gc_manager(struct f2fs_sb_info *sbi); int f2fs_gc_range(struct f2fs_sb_info *sbi, unsigned int start_seg, unsigned int end_seg, bool dry_run, unsigned int dry_run_sections); +void f2fs_reset_gc_victim_resource(struct f2fs_sb_info *sbi, + unsigned int start, unsigned int end); int f2fs_resize_fs(struct file *filp, __u64 block_count); int __init f2fs_create_garbage_collection_cache(void); void f2fs_destroy_garbage_collection_cache(void); diff --git a/fs/f2fs/file.c b/fs/f2fs/file.c index 4b52c56d71f0..642f06f2faea 100644 --- a/fs/f2fs/file.c +++ b/fs/f2fs/file.c @@ -813,13 +813,19 @@ int f2fs_do_truncate_blocks(struct inode *inode, u64 = from, bool lock) =20 if (IS_DEVICE_ALIASING(inode)) { struct extent_tree *et =3D F2FS_I(inode)->extent_tree[EX_READ]; - struct extent_info ei =3D et->largest; + struct extent_info ei; + + read_lock(&et->lock); + ei =3D et->largest; + read_unlock(&et->lock); =20 f2fs_invalidate_blocks(sbi, ei.blk, ei.len); =20 dec_valid_block_count(sbi, inode, ei.len); f2fs_update_time(sbi, REQ_TIME); =20 + f2fs_drop_extent_tree(inode); + f2fs_folio_put(ifolio, true); goto out; } @@ -1100,8 +1106,9 @@ int f2fs_setattr(struct mnt_idmap *idmap, struct dent= ry *dentry, if ((attr->ia_valid & ATTR_SIZE)) { if (mapping_large_folio_support(inode->i_mapping)) return -EOPNOTSUPP; - if (!f2fs_is_compress_backend_ready(inode) || - IS_DEVICE_ALIASING(inode)) + if (IS_DEVICE_ALIASING(inode)) + return -EPERM; + if (!f2fs_is_compress_backend_ready(inode)) return -EOPNOTSUPP; if (is_inode_flag_set(inode, FI_COMPRESS_RELEASED) && !IS_ALIGNED(attr->ia_size, @@ -2130,6 +2137,9 @@ static int f2fs_setflags_common(struct inode *inode, = u32 iflags, u32 mask) if (IS_NOQUOTA(inode)) return -EPERM; =20 + if (IS_DEVICE_ALIASING(inode)) + return -EPERM; + if ((iflags ^ masked_flags) & F2FS_CASEFOLD_FL) { if (!f2fs_sb_has_casefold(F2FS_I_SB(inode))) return -EOPNOTSUPP; @@ -2678,6 +2688,17 @@ static int f2fs_ioc_get_encryption_policy(struct fil= e *filp, unsigned long arg) return fscrypt_ioctl_get_policy(filp, (void __user *)arg); } =20 +static int f2fs_ioc_get_dev_alias_status(struct file *filp, unsigned long = arg) +{ + struct inode *inode =3D file_inode(filp); + + if (!IS_DEVICE_ALIASING(inode)) + return -EINVAL; + + return put_user(F2FS_HAS_BLOCKS(inode) ? F2FS_DEV_ALIAS_STATUS_RESERVED : + F2FS_DEV_ALIAS_STATUS_RELEASED, (u32 __user *)arg); +} + static int f2fs_ioc_get_encryption_pwsalt(struct file *filp, unsigned long= arg) { struct inode *inode =3D file_inode(filp); @@ -3616,6 +3637,248 @@ static int f2fs_ioc_get_dev_alias_file(struct file = *filp, unsigned long arg) (u32 __user *)arg); } =20 +static bool f2fs_get_dev_alias_extent(struct f2fs_sb_info *sbi, + struct dentry *dentry, + struct extent_info *ei) +{ + int i; + + for (i =3D 1; i < sbi->s_ndevs; i++) { + char *name =3D strrchr(FDEV(i).path, '/'); + + name =3D name ? name + 1 : FDEV(i).path; + if (!strcmp(name, dentry->d_name.name)) { + ei->blk =3D FDEV(i).start_blk; + ei->len =3D FDEV(i).total_segments << sbi->log_blocks_per_seg; + ei->fofs =3D 0; + return true; + } + } + return false; +} + +static int f2fs_ioc_reserve_dev_alias(struct file *filp) +{ + struct inode *inode =3D file_inode(filp); + struct f2fs_sb_info *sbi =3D F2FS_I_SB(inode); + struct extent_tree *et =3D F2FS_I(inode)->extent_tree[EX_READ]; + struct extent_info ei; + struct cp_control cpc =3D { CP_SYNC, 0, 0, 0 }; + struct f2fs_lock_context lc, glc; + blkcnt_t count; + unsigned int start, end; + int type, err; + + if (!capable(CAP_SYS_ADMIN)) + return -EPERM; + + if (unlikely(is_sbi_flag_set(sbi, SBI_CP_DISABLED))) + return -EINVAL; + + err =3D mnt_want_write_file(filp); + if (err) + return err; + + inode_lock(inode); + + if (!IS_DEVICE_ALIASING(inode)) { + err =3D -EINVAL; + goto out_inode_unlock; + } + + if (F2FS_HAS_BLOCKS(inode)) { + err =3D 0; + goto out_inode_unlock; + } + + if (!f2fs_get_dev_alias_extent(sbi, filp->f_path.dentry, &ei)) { + f2fs_warn(sbi, "device alias file (%s, ino=3D%llu) has no matching devic= e", + filp->f_path.dentry->d_name.name, + (unsigned long long)inode->i_ino); + set_sbi_flag(sbi, SBI_NEED_FSCK); + f2fs_handle_error(sbi, ERROR_CORRUPTED_INODE); + err =3D -EFSCORRUPTED; + goto out_inode_unlock; + } + + spin_lock(&sbi->stat_lock); + if (sbi->total_valid_block_count + ei.len > + get_available_block_count(sbi, inode, true)) { + spin_unlock(&sbi->stat_lock); + err =3D -ENOSPC; + goto out_inode_unlock; + } + sbi->alias_reserved_blocks +=3D ei.len; + spin_unlock(&sbi->stat_lock); + + spin_lock(&FREE_I(sbi)->segmap_lock); + FDEV(f2fs_target_device_index(sbi, ei.blk)).is_reserving =3D true; + spin_unlock(&FREE_I(sbi)->segmap_lock); + + start =3D GET_SEGNO(sbi, ei.blk); + end =3D GET_SEGNO(sbi, ei.blk + ei.len - 1); + + /* Acquire gc_lock for victim reset, curseg resize, and range GC */ + f2fs_down_write_trace(&sbi->gc_lock, &glc); + + /* Reset the victim information to prevent GC from targeting the range */ + f2fs_reset_gc_victim_resource(sbi, start, end); + + /* Move out cursegs from the target range */ + for (type =3D CURSEG_HOT_DATA; type < NR_CURSEG_PERSIST_TYPE; type++) { + err =3D f2fs_allocate_segment_for_resize(sbi, type, start, end); + if (err) + goto out_gc_unlock; + } + + f2fs_lock_op(sbi, &lc); + + if (unlikely(is_sbi_flag_set(sbi, SBI_CP_DISABLED))) { + err =3D -EINVAL; + f2fs_unlock_op(sbi, &lc); + goto out_gc_unlock; + } + + /* do GC to move out valid blocks in the range all at once! */ + err =3D f2fs_gc_range(sbi, start, end, false, 0); + if (err) { + f2fs_unlock_op(sbi, &lc); + goto out_gc_unlock; + } + + count =3D ei.len; + err =3D inc_valid_block_count(sbi, inode, &count, false); + if (err) { + f2fs_unlock_op(sbi, &lc); + goto out_gc_unlock; + } + + spin_lock(&sbi->stat_lock); + sbi->alias_reserved_blocks -=3D ei.len; + spin_unlock(&sbi->stat_lock); + + if (et) { + write_lock(&et->lock); + et->largest =3D ei; + write_unlock(&et->lock); + } + clear_inode_flag(inode, FI_NO_EXTENT); + + f2fs_reserve_device_alias(sbi, ei.blk, ei.len); + + i_size_write(inode, (loff_t)ei.len << PAGE_SHIFT); + f2fs_update_inode_page(inode); + + spin_lock(&FREE_I(sbi)->segmap_lock); + FDEV(f2fs_target_device_index(sbi, ei.blk)).is_reserving =3D false; + spin_unlock(&FREE_I(sbi)->segmap_lock); + + f2fs_unlock_op(sbi, &lc); + f2fs_up_write_trace(&sbi->gc_lock, &glc); + + inode_unlock(inode); + mnt_drop_write_file(filp); + + err =3D f2fs_write_checkpoint(sbi, &cpc); + return err; + +out_gc_unlock: + spin_lock(&sbi->stat_lock); + sbi->alias_reserved_blocks -=3D ei.len; + spin_unlock(&sbi->stat_lock); + + spin_lock(&FREE_I(sbi)->segmap_lock); + FDEV(f2fs_target_device_index(sbi, ei.blk)).is_reserving =3D false; + spin_unlock(&FREE_I(sbi)->segmap_lock); + f2fs_up_write_trace(&sbi->gc_lock, &glc); + + inode_unlock(inode); + mnt_drop_write_file(filp); + return err; + +out_inode_unlock: + inode_unlock(inode); + mnt_drop_write_file(filp); + return err; +} + +static int f2fs_ioc_release_dev_alias(struct file *filp) +{ + struct inode *inode =3D file_inode(filp); + struct f2fs_sb_info *sbi =3D F2FS_I_SB(inode); + struct extent_tree *et =3D F2FS_I(inode)->extent_tree[EX_READ]; + struct extent_info ei =3D {0, }; + struct cp_control cpc =3D { CP_SYNC, 0, 0, 0 }; + struct f2fs_lock_context lc, glc; + int err; + + if (!capable(CAP_SYS_ADMIN)) + return -EPERM; + + if (unlikely(is_sbi_flag_set(sbi, SBI_CP_DISABLED))) + return -EINVAL; + + err =3D mnt_want_write_file(filp); + if (err) + return err; + + inode_lock(inode); + + if (!IS_DEVICE_ALIASING(inode)) { + err =3D -EINVAL; + goto out_inode_unlock; + } + + if (!F2FS_HAS_BLOCKS(inode)) { + err =3D 0; + goto out_inode_unlock; + } + + err =3D filemap_write_and_wait(inode->i_mapping); + if (err) + goto out_inode_unlock; + + read_lock(&et->lock); + ei =3D et->largest; + read_unlock(&et->lock); + + f2fs_down_write_trace(&sbi->gc_lock, &glc); + f2fs_lock_op(sbi, &lc); + + if (unlikely(is_sbi_flag_set(sbi, SBI_CP_DISABLED))) { + err =3D -EINVAL; + f2fs_unlock_op(sbi, &lc); + f2fs_up_write_trace(&sbi->gc_lock, &glc); + goto out_inode_unlock; + } + + truncate_setsize(inode, 0); + + err =3D f2fs_truncate_blocks(inode, 0, false); + if (err) { + i_size_write(inode, (loff_t)ei.len << PAGE_SHIFT); + f2fs_unlock_op(sbi, &lc); + f2fs_up_write_trace(&sbi->gc_lock, &glc); + goto out_inode_unlock; + } + + f2fs_update_inode_page(inode); + + f2fs_unlock_op(sbi, &lc); + f2fs_up_write_trace(&sbi->gc_lock, &glc); + + inode_unlock(inode); + mnt_drop_write_file(filp); + + err =3D f2fs_write_checkpoint(sbi, &cpc); + return err; + +out_inode_unlock: + inode_unlock(inode); + mnt_drop_write_file(filp); + return err; +} + static int f2fs_ioc_io_prio(struct file *filp, unsigned long arg) { struct inode *inode =3D file_inode(filp); @@ -4742,8 +5005,14 @@ static long __f2fs_ioctl(struct file *filp, unsigned= int cmd, unsigned long arg) return f2fs_ioc_compress_file(filp); case F2FS_IOC_GET_DEV_ALIAS_FILE: return f2fs_ioc_get_dev_alias_file(filp, arg); + case F2FS_IOC_GET_DEV_ALIAS_STATUS: + return f2fs_ioc_get_dev_alias_status(filp, arg); case F2FS_IOC_IO_PRIO: return f2fs_ioc_io_prio(filp, arg); + case F2FS_IOC_RESERVE_DEV_ALIAS: + return f2fs_ioc_reserve_dev_alias(filp); + case F2FS_IOC_RELEASE_DEV_ALIAS: + return f2fs_ioc_release_dev_alias(filp); default: return -ENOTTY; } @@ -5530,7 +5799,10 @@ long f2fs_compat_ioctl(struct file *file, unsigned i= nt cmd, unsigned long arg) case F2FS_IOC_DECOMPRESS_FILE: case F2FS_IOC_COMPRESS_FILE: case F2FS_IOC_GET_DEV_ALIAS_FILE: + case F2FS_IOC_GET_DEV_ALIAS_STATUS: case F2FS_IOC_IO_PRIO: + case F2FS_IOC_RESERVE_DEV_ALIAS: + case F2FS_IOC_RELEASE_DEV_ALIAS: break; default: return -ENOIOCTLCMD; diff --git a/fs/f2fs/gc.c b/fs/f2fs/gc.c index ffaa7ba76a1b..93bcb35a5b5d 100644 --- a/fs/f2fs/gc.c +++ b/fs/f2fs/gc.c @@ -2197,29 +2197,37 @@ int f2fs_gc_range(struct f2fs_sb_info *sbi, return 0; } =20 +void f2fs_reset_gc_victim_resource(struct f2fs_sb_info *sbi, + unsigned int start, unsigned int end) +{ + int i; + + mutex_lock(&DIRTY_I(sbi)->seglist_lock); + for (i =3D 0; i < MAX_GC_POLICY; i++) + if (SIT_I(sbi)->last_victim[i] >=3D start && + SIT_I(sbi)->last_victim[i] <=3D end) + SIT_I(sbi)->last_victim[i] =3D 0; + + for (i =3D BG_GC; i <=3D FG_GC; i++) + if (sbi->next_victim_seg[i] >=3D start && + sbi->next_victim_seg[i] <=3D end) + sbi->next_victim_seg[i] =3D NULL_SEGNO; + mutex_unlock(&DIRTY_I(sbi)->seglist_lock); +} + static int free_segment_range(struct f2fs_sb_info *sbi, unsigned int secs, bool dry_run) { unsigned int next_inuse, start, end; struct cp_control cpc =3D { CP_RESIZE, 0, 0, 0 }; - int gc_mode, gc_type; int err =3D 0; int type; =20 - /* Force block allocation for GC */ MAIN_SECS(sbi) -=3D secs; start =3D MAIN_SECS(sbi) * SEGS_PER_SEC(sbi); end =3D MAIN_SEGS(sbi) - 1; =20 - mutex_lock(&DIRTY_I(sbi)->seglist_lock); - for (gc_mode =3D 0; gc_mode < MAX_GC_POLICY; gc_mode++) - if (SIT_I(sbi)->last_victim[gc_mode] >=3D start) - SIT_I(sbi)->last_victim[gc_mode] =3D 0; - - for (gc_type =3D BG_GC; gc_type <=3D FG_GC; gc_type++) - if (sbi->next_victim_seg[gc_type] >=3D start) - sbi->next_victim_seg[gc_type] =3D NULL_SEGNO; - mutex_unlock(&DIRTY_I(sbi)->seglist_lock); + f2fs_reset_gc_victim_resource(sbi, start, end); =20 /* Move out cursegs from the target range */ for (type =3D CURSEG_HOT_DATA; type < NR_CURSEG_PERSIST_TYPE; type++) { diff --git a/fs/f2fs/namei.c b/fs/f2fs/namei.c index cac03b8e91a1..8c3b57987f6c 100644 --- a/fs/f2fs/namei.c +++ b/fs/f2fs/namei.c @@ -425,6 +425,9 @@ static int f2fs_link(struct dentry *old_dentry, struct = inode *dir, if (!f2fs_is_checkpoint_ready(sbi)) return -ENOSPC; =20 + if (IS_DEVICE_ALIASING(inode)) + return -EPERM; + err =3D fscrypt_prepare_link(old_dentry, dir, dentry); if (err) return err; @@ -568,6 +571,9 @@ static int f2fs_unlink(struct inode *dir, struct dentry= *dentry) =20 trace_f2fs_unlink_enter(dir, dentry); =20 + if (IS_DEVICE_ALIASING(inode)) + return -EPERM; + if (unlikely(f2fs_cp_error(sbi))) { err =3D -EIO; goto out; @@ -946,6 +952,9 @@ static int f2fs_rename(struct mnt_idmap *idmap, struct = inode *old_dir, bool old_is_dir =3D S_ISDIR(old_inode->i_mode); int err; =20 + if (IS_DEVICE_ALIASING(old_inode)) + return -EPERM; + if (unlikely(f2fs_cp_error(sbi))) return -EIO; if (!f2fs_is_checkpoint_ready(sbi)) @@ -1016,6 +1025,8 @@ static int f2fs_rename(struct mnt_idmap *idmap, struc= t inode *old_dir, } =20 if (new_inode) { + if (IS_DEVICE_ALIASING(new_inode)) + return -EPERM; =20 err =3D -ENOTEMPTY; if (old_is_dir && !f2fs_empty_dir(new_inode)) @@ -1143,6 +1154,9 @@ static int f2fs_cross_rename(struct inode *old_dir, s= truct dentry *old_dentry, int old_nlink =3D 0, new_nlink =3D 0; int err; =20 + if (IS_DEVICE_ALIASING(old_inode) || IS_DEVICE_ALIASING(new_inode)) + return -EPERM; + if (unlikely(f2fs_cp_error(sbi))) return -EIO; if (!f2fs_is_checkpoint_ready(sbi)) diff --git a/fs/f2fs/segment.c b/fs/f2fs/segment.c index d71ddb3ee918..1202bde3390a 100644 --- a/fs/f2fs/segment.c +++ b/fs/f2fs/segment.c @@ -2502,35 +2502,42 @@ static int update_sit_entry_for_alloc(struct f2fs_s= b_info *sbi, struct seg_entry unsigned int segno, block_t blkaddr, unsigned int offset, int del) { bool exist; + int del_count =3D del; + int i; =20 - exist =3D f2fs_test_and_set_bit(offset, se->cur_valid_map); - if (unlikely(exist)) { - f2fs_err(sbi, "Bitmap was wrongly set, blk:%u", blkaddr); - f2fs_bug_on(sbi, 1); - se->valid_blocks--; - del =3D 0; - } + f2fs_bug_on(sbi, GET_SEGNO(sbi, blkaddr) !=3D GET_SEGNO(sbi, blkaddr + de= l_count - 1)); =20 - if (f2fs_block_unit_discard(sbi) && - !f2fs_test_and_set_bit(offset, se->discard_map)) - sbi->discard_blks--; + for (i =3D 0; i < del_count; i++) { + exist =3D f2fs_test_and_set_bit(offset + i, se->cur_valid_map); + if (unlikely(exist)) { + f2fs_err(sbi, "Bitmap was wrongly set, blk:%u", blkaddr + i); + f2fs_bug_on(sbi, 1); + se->valid_blocks--; + del -=3D 1; + continue; + } =20 - /* - * SSR should never reuse block which is checkpointed - * or newly invalidated. - */ - if (!is_sbi_flag_set(sbi, SBI_CP_DISABLED)) { - if (!f2fs_test_and_set_bit(offset, se->ckpt_valid_map)) { - se->ckpt_valid_blocks++; - if (__is_large_section(sbi)) - get_sec_entry(sbi, segno)->ckpt_valid_blocks++; + if (f2fs_block_unit_discard(sbi) && + !f2fs_test_and_set_bit(offset + i, se->discard_map)) + sbi->discard_blks--; + + /* + * SSR should never reuse block which is checkpointed + * or newly invalidated. + */ + if (!is_sbi_flag_set(sbi, SBI_CP_DISABLED)) { + if (!f2fs_test_and_set_bit(offset + i, se->ckpt_valid_map)) { + se->ckpt_valid_blocks++; + if (__is_large_section(sbi)) + get_sec_entry(sbi, segno)->ckpt_valid_blocks++; + } } - } =20 - if (!f2fs_test_bit(offset, se->ckpt_valid_map)) { - se->ckpt_valid_blocks +=3D del; - if (__is_large_section(sbi)) - get_sec_entry(sbi, segno)->ckpt_valid_blocks +=3D del; + if (!f2fs_test_bit(offset + i, se->ckpt_valid_map)) { + se->ckpt_valid_blocks +=3D 1; + if (__is_large_section(sbi)) + get_sec_entry(sbi, segno)->ckpt_valid_blocks +=3D 1; + } } =20 if (__is_large_section(sbi)) @@ -2585,9 +2592,14 @@ void f2fs_invalidate_blocks(struct f2fs_sb_info *sbi= , block_t addr, unsigned int segno =3D GET_SEGNO(sbi, addr); struct sit_info *sit_i =3D SIT_I(sbi); block_t addr_start =3D addr, addr_end =3D addr + len - 1; - unsigned int seg_num =3D GET_SEGNO(sbi, addr_end) - segno + 1; + unsigned int seg_num; unsigned int i =3D 1, max_blocks =3D sbi->blocks_per_seg, cnt; =20 + if (len =3D=3D 0) + return; + + seg_num =3D GET_SEGNO(sbi, addr_end) - segno + 1; + f2fs_bug_on(sbi, addr =3D=3D NULL_ADDR); if (addr =3D=3D NEW_ADDR || addr =3D=3D COMPRESS_ADDR) return; @@ -2620,6 +2632,52 @@ void f2fs_invalidate_blocks(struct f2fs_sb_info *sbi= , block_t addr, up_write(&sit_i->sentry_lock); } =20 +void f2fs_reserve_device_alias(struct f2fs_sb_info *sbi, block_t addr, + unsigned int len) +{ + unsigned int segno =3D GET_SEGNO(sbi, addr); + struct sit_info *sit_i =3D SIT_I(sbi); + block_t addr_start =3D addr, addr_end =3D addr + len - 1; + unsigned int seg_num; + unsigned int i =3D 1, max_blocks =3D sbi->blocks_per_seg, cnt; + + if (len =3D=3D 0) + return; + + seg_num =3D GET_SEGNO(sbi, addr_end) - segno + 1; + + down_write(&sit_i->sentry_lock); + + if (seg_num =3D=3D 1) + cnt =3D len; + else + cnt =3D max_blocks - GET_BLKOFF_FROM_SEG0(sbi, addr); + + do { + update_segment_mtime(sbi, addr_start, 0); + update_sit_entry(sbi, addr_start, cnt); + __set_test_and_inuse(sbi, segno); + + /* Remove the segment from PRE (prefree) to prevent checkpoint from free= ing it! */ + mutex_lock(&DIRTY_I(sbi)->seglist_lock); + if (test_and_clear_bit(segno, DIRTY_I(sbi)->dirty_segmap[PRE])) + DIRTY_I(sbi)->nr_dirty[PRE]--; + mutex_unlock(&DIRTY_I(sbi)->seglist_lock); + + /* add it into dirty seglist */ + locate_dirty_segment(sbi, segno); + + /* update @addr_start and @cnt and @segno */ + addr_start =3D START_BLOCK(sbi, ++segno); + if (++i =3D=3D seg_num) + cnt =3D GET_BLKOFF_FROM_SEG0(sbi, addr_end) + 1; + else + cnt =3D max_blocks; + } while (i <=3D seg_num); + + up_write(&sit_i->sentry_lock); +} + bool f2fs_is_checkpointed_data(struct f2fs_sb_info *sbi, block_t blkaddr) { struct sit_info *sit_i =3D SIT_I(sbi); @@ -2758,8 +2816,13 @@ static int is_next_segment_free(struct f2fs_sb_info = *sbi, unsigned int segno =3D curseg->segno + 1; struct free_segmap_info *free_i =3D FREE_I(sbi); =20 - if (segno < MAIN_SEGS(sbi) && segno % SEGS_PER_SEC(sbi)) + if (segno < MAIN_SEGS(sbi) && segno % SEGS_PER_SEC(sbi)) { + int devi =3D f2fs_target_device_index(sbi, START_BLOCK(sbi, segno)); + + if (f2fs_dev_is_reserving(sbi, devi)) + return 0; return !test_bit(segno, free_i->free_segmap); + } return 0; } =20 @@ -2778,7 +2841,8 @@ static int get_new_segment(struct f2fs_sb_info *sbi, unsigned int alloc_policy =3D sbi->allocate_section_policy; unsigned int alloc_hint =3D sbi->allocate_section_hint; bool init =3D true; - int i; + bool looped =3D false; + int i, devi; int ret =3D 0; =20 spin_lock(&free_i->segmap_lock); @@ -2791,8 +2855,13 @@ static int get_new_segment(struct f2fs_sb_info *sbi, if (!new_sec && ((*newseg + 1) % SEGS_PER_SEC(sbi))) { segno =3D find_next_zero_bit(free_i->free_segmap, GET_SEG_FROM_SEC(sbi, hint + 1), *newseg + 1); - if (segno < GET_SEG_FROM_SEC(sbi, hint + 1)) + if (segno < GET_SEG_FROM_SEC(sbi, hint + 1)) { + devi =3D f2fs_target_device_index(sbi, START_BLOCK(sbi, segno)); + + if (f2fs_dev_is_alloc_blocked(sbi, devi, pinning)) + goto find_other_zone; goto got_it; + } } =20 #ifdef CONFIG_BLK_DEV_ZONED @@ -2828,33 +2897,42 @@ static int get_new_segment(struct f2fs_sb_info *sbi, find_other_zone: secno =3D find_next_zero_bit(free_i->free_secmap, MAIN_SECS(sbi), hint); =20 -#ifdef CONFIG_BLK_DEV_ZONED - if (secno >=3D MAIN_SECS(sbi) && f2fs_sb_has_blkzoned(sbi)) { - /* Write only to sequential zones */ - if (sbi->blkzone_alloc_policy =3D=3D BLKZONE_ALLOC_ONLY_SEQ) { - hint =3D GET_SEC_FROM_SEG(sbi, sbi->first_seq_zone_segno); - secno =3D find_next_zero_bit(free_i->free_secmap, MAIN_SECS(sbi), hint); - } else - secno =3D find_first_zero_bit(free_i->free_secmap, - MAIN_SECS(sbi)); - if (secno >=3D MAIN_SECS(sbi)) { - ret =3D -ENOSPC; - f2fs_bug_on(sbi, 1); - goto out_unlock; - } - } -#endif - if (secno >=3D MAIN_SECS(sbi)) { - secno =3D find_first_zero_bit(free_i->free_secmap, - MAIN_SECS(sbi)); - if (secno >=3D MAIN_SECS(sbi)) { + if (looped) { ret =3D -ENOSPC; f2fs_bug_on(sbi, !pinning); goto out_unlock; } + hint =3D 0; +#ifdef CONFIG_BLK_DEV_ZONED + /* Write only to sequential zones */ + if (f2fs_sb_has_blkzoned(sbi) && + sbi->blkzone_alloc_policy =3D=3D BLKZONE_ALLOC_ONLY_SEQ) + hint =3D GET_SEC_FROM_SEG(sbi, sbi->first_seq_zone_segno); +#endif + looped =3D true; + goto find_other_zone; } + segno =3D GET_SEG_FROM_SEC(sbi, secno); + + devi =3D f2fs_target_device_index(sbi, START_BLOCK(sbi, segno)); + + if (f2fs_dev_is_alloc_blocked(sbi, devi, pinning)) { + while (devi < sbi->s_ndevs && + f2fs_dev_is_alloc_blocked(sbi, devi, pinning)) { + unsigned int end_segno =3D GET_SEGNO(sbi, FDEV(devi).end_blk); + + hint =3D GET_SEC_FROM_SEG(sbi, end_segno) + 1; + devi++; + } + goto find_other_zone; + } + + if (sec_usage_check(sbi, secno)) { + hint =3D secno + 1; + goto find_other_zone; + } zoneno =3D GET_ZONE_FROM_SEC(sbi, secno); =20 /* give up on finding another zone */ diff --git a/fs/f2fs/segment.h b/fs/f2fs/segment.h index b0c06b3580b4..107508ae7eb3 100644 --- a/fs/f2fs/segment.h +++ b/fs/f2fs/segment.h @@ -954,10 +954,32 @@ static inline block_t sum_blk_addr(struct f2fs_sb_inf= o *sbi, int base, int type) - (base + 1) + type; } =20 +static inline bool f2fs_dev_is_reserving(struct f2fs_sb_info *sbi, int dev= i) +{ + if (!f2fs_sb_has_device_alias(sbi)) + return false; + return FDEV(devi).is_reserving; +} + +static inline bool f2fs_dev_is_alloc_blocked(struct f2fs_sb_info *sbi, + int devi, bool pinning) +{ + if (!f2fs_sb_has_device_alias(sbi)) + return false; + return (pinning && FDEV(devi).has_alias) || FDEV(devi).is_reserving; +} + static inline bool sec_usage_check(struct f2fs_sb_info *sbi, unsigned int = secno) { if (is_cursec(sbi, secno) || (sbi->cur_victim_sec =3D=3D secno)) return true; + if (f2fs_sb_has_device_alias(sbi)) { + block_t start_blk =3D START_BLOCK(sbi, GET_SEG_FROM_SEC(sbi, secno)); + int devi =3D f2fs_target_device_index(sbi, start_blk); + + if (f2fs_dev_is_reserving(sbi, devi)) + return true; + } return false; } =20 diff --git a/fs/f2fs/super.c b/fs/f2fs/super.c index c448d992ff2a..bca1475cb573 100644 --- a/fs/f2fs/super.c +++ b/fs/f2fs/super.c @@ -5001,6 +5001,38 @@ static void f2fs_tuning_parameters(struct f2fs_sb_in= fo *sbi) sbi->readdir_ra =3D true; } =20 +static void f2fs_restore_device_alias(struct f2fs_sb_info *sbi) +{ + struct inode *root =3D d_inode(sbi->sb->s_root); + struct f2fs_dir_entry *de; + struct folio *folio; + int i; + + if (!f2fs_sb_has_device_alias(sbi)) + return; + + for (i =3D 1; i < sbi->s_ndevs; i++) { + char *name =3D strrchr(FDEV(i).path, '/'); + struct qstr qstr; + + name =3D name ? name + 1 : FDEV(i).path; + qstr.name =3D name; + qstr.len =3D strlen(name); + + de =3D f2fs_find_entry(root, &qstr, &folio); + if (de) { + struct inode *inode =3D f2fs_iget(sbi->sb, le32_to_cpu(de->ino)); + + if (!IS_ERR(inode)) { + if (IS_DEVICE_ALIASING(inode)) + FDEV(i).has_alias =3D true; + iput(inode); + } + f2fs_folio_put(folio, 0); + } + } +} + static int f2fs_fill_super(struct super_block *sb, struct fs_context *fc) { struct f2fs_fs_context *ctx =3D fc->fs_private; @@ -5209,6 +5241,7 @@ static int f2fs_fill_super(struct super_block *sb, st= ruct fs_context *fc) sbi->last_valid_block_count =3D sbi->total_valid_block_count; sbi->reserved_blocks =3D 0; sbi->current_reserved_blocks =3D 0; + sbi->alias_reserved_blocks =3D 0; limit_reserve_root(sbi); adjust_unusable_cap_perc(sbi); =20 @@ -5436,6 +5469,8 @@ static int f2fs_fill_super(struct super_block *sb, st= ruct fs_context *fc) f2fs_update_time(sbi, REQ_TIME); clear_sbi_flag(sbi, SBI_CP_DISABLED_QUICK); =20 + f2fs_restore_device_alias(sbi); + sbi->umount_lock_holder =3D NULL; return 0; =20 diff --git a/include/uapi/linux/f2fs.h b/include/uapi/linux/f2fs.h index 795e26258355..4409ada2fecb 100644 --- a/include/uapi/linux/f2fs.h +++ b/include/uapi/linux/f2fs.h @@ -45,6 +45,9 @@ #define F2FS_IOC_START_ATOMIC_REPLACE _IO(F2FS_IOCTL_MAGIC, 25) #define F2FS_IOC_GET_DEV_ALIAS_FILE _IOR(F2FS_IOCTL_MAGIC, 26, __u32) #define F2FS_IOC_IO_PRIO _IOW(F2FS_IOCTL_MAGIC, 27, __u32) +#define F2FS_IOC_RESERVE_DEV_ALIAS _IO(F2FS_IOCTL_MAGIC, 28) +#define F2FS_IOC_RELEASE_DEV_ALIAS _IO(F2FS_IOCTL_MAGIC, 29) +#define F2FS_IOC_GET_DEV_ALIAS_STATUS _IOR(F2FS_IOCTL_MAGIC, 30, __u32) =20 /* * should be same as XFS_IOC_GOINGDOWN. @@ -70,6 +73,10 @@ enum { F2FS_IOPRIO_MAX, }; =20 +/* for F2FS_IOC_GET_DEV_ALIAS_STATUS */ +#define F2FS_DEV_ALIAS_STATUS_RELEASED 0 +#define F2FS_DEV_ALIAS_STATUS_RESERVED 1 + struct f2fs_gc_range { __u32 sync; __u64 start; --=20 2.55.0.229.g6434b31f56-goog