From nobody Sat Oct 3 04:22:53 2026 Received: from mx0b-001b2d01.pphosted.com (mx0b-001b2d01.pphosted.com [148.163.158.5]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 1141833290F; Wed, 5 Aug 2026 06:34:03 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=148.163.158.5 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785911645; cv=none; b=iKwHrxLBKMAEjeicFgc3T9OO2xfpDQeZgKPubwJjQG1kyifRREOh5W5N1euyXXahkwCReQKZONayJNi7ZIz+cy+1uJMHuhC80V5RXsoKvWTX5CKfNLBlcduBZlG+GeCGDAFv55P4hCt4BdufAQ0qUxbtr2RLnnSTAMO80EUs5+4= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785911645; c=relaxed/simple; bh=flfKAXGK9Mx9z//gL69lq+sz4HqSHGpjSgJGHBE12cY=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=ZgOYtj0EG23jbDK6E4T1Y86+1w59/zSKikdzgENDcA2YMtg5ezH6ktYn7gore2/MxJW41runJTP1G1b9hI46kaE3eoZCoaWlMOw26Lv8BvFPsWvMyMMUTVXZgK9dNhV7WCURNxC6vowE1k3NvuMiamJmZ7V2N3sI66kvdj3T/BQ= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.ibm.com; spf=pass smtp.mailfrom=linux.ibm.com; dkim=pass (2048-bit key) header.d=ibm.com header.i=@ibm.com header.b=iBtLqhFB; arc=none smtp.client-ip=148.163.158.5 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.ibm.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.ibm.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=ibm.com header.i=@ibm.com header.b="iBtLqhFB" Received: from pps.filterd (m0353725.ppops.net [127.0.0.1]) by mx0a-001b2d01.pphosted.com (8.18.1.11/8.18.1.11) with ESMTP id 6755mYMR2910700; Wed, 5 Aug 2026 06:28:31 GMT DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=ibm.com; h=cc :content-transfer-encoding:date:from:in-reply-to:message-id :mime-version:references:subject:to; s=pp1; bh=7mkcadJVmAqJI9bua p6E5CxELlfihDwaquThutvKHd4=; b=iBtLqhFBgUmsk9kJLMFftzKOe75YexqsD YXuBsgZi9G5Qgm1zsshk5BcyA7kZMjYOJM9nI5MLQLeLYaJz+/xJUMyuCmBdDs+H xNuhNcptmzRo0Luoa7Bp3DFUpXVyQ0QQ0mCAdsy+TECeBLTUV0I024g8vXATeiL3 IOKw9J6KiBRWsQ23xMZx0G4sUKsjQxMit/xFu2SAIDObi9fevvD926v/aVyNSrXT C7MX19drUWgihVc6G8BYBWVAQGS9BbUYS+B9lH3uscO5KXTQnZcRSCxoj2V8Pym0 xQfWSo1U6QYS2dLXDuMDFfTJwU9+vdGc76x5euZ7tgEqMls0WOOkQ== Received: from ppma13.dal12v.mail.ibm.com (dd.9e.1632.ip4.static.sl-reverse.com [50.22.158.221]) by mx0a-001b2d01.pphosted.com (PPS) with ESMTPS id 4fs77g99ka-1 (version=TLSv1.2 cipher=ECDHE-RSA-AES256-GCM-SHA384 bits=256 verify=NOT); Wed, 05 Aug 2026 06:28:30 +0000 (GMT) Received: from pps.filterd (ppma13.dal12v.mail.ibm.com [127.0.0.1]) by ppma13.dal12v.mail.ibm.com (8.18.1.7/8.18.1.7) with ESMTP id 6756QJdg026819; Wed, 5 Aug 2026 06:28:30 GMT Received: from smtprelay01.fra02v.mail.ibm.com ([9.218.2.227]) by ppma13.dal12v.mail.ibm.com (PPS) with ESMTPS id 4fswbgd4ja-1 (version=TLSv1.2 cipher=ECDHE-RSA-AES256-GCM-SHA384 bits=256 verify=NOT); Wed, 05 Aug 2026 06:28:30 +0000 (GMT) Received: from smtpav07.fra02v.mail.ibm.com (smtpav07.fra02v.mail.ibm.com [10.20.54.106]) by smtprelay01.fra02v.mail.ibm.com (8.14.9/8.14.9/NCO v10.0) with ESMTP id 6756SSBS37224742 (version=TLSv1/SSLv3 cipher=DHE-RSA-AES256-GCM-SHA384 bits=256 verify=OK); Wed, 5 Aug 2026 06:28:28 GMT Received: from smtpav07.fra02v.mail.ibm.com (unknown [127.0.0.1]) by IMSVA (Postfix) with ESMTP id 37DD720043; Wed, 5 Aug 2026 06:28:28 +0000 (GMT) Received: from smtpav07.fra02v.mail.ibm.com (unknown [127.0.0.1]) by IMSVA (Postfix) with ESMTP id 0F80020040; Wed, 5 Aug 2026 06:28:24 +0000 (GMT) Received: from li-dc0c254c-257c-11b2-a85c-98b6c1322444.ibm.com (unknown [9.124.211.239]) by smtpav07.fra02v.mail.ibm.com (Postfix) with ESMTP; Wed, 5 Aug 2026 06:28:23 +0000 (GMT) From: Ojaswin Mujoo To: Christian Brauner , linux-fsdevel@vger.kernel.org Cc: "Darrick J . Wong" , Carlos Maiolino , Alexander Viro , Jan Kara , Matthew Wilcox , Andrew Morton , Ritesh Harjani , Zhang Yi , Christoph Hellwig , Dave Chinner , Daniel Gomez , Pankaj Raghav , Theodore Tso , linux-xfs@vger.kernel.org, linux-kernel@vger.kernel.org, linux-mm@kvack.org Subject: [RFC PATCH v3 01/11] fs: Add counter to track inflight writes that need stable pages Date: Wed, 5 Aug 2026 11:58:07 +0530 Message-ID: <548c55490029490e68287aeae3d737454fa78e0e.1785908600.git.ojaswin@linux.ibm.com> X-Mailer: git-send-email 2.55.0 In-Reply-To: References: Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-TM-AS-GCONF: 00 X-Proofpoint-Reinject: loops=2 maxloops=12 X-Proofpoint-Spam-Details-Enc: AW1haW4tMjYwODA1MDA0NyBTYWx0ZWRfX15tpIh4spCjU AUcBnujClBo+DcVOKF9/a4d+yMeUkxJ5qnaJaksPvwrcCxrQG2PrxUfOZkSb175hfmMH5EDv3mZ 6TB0NfHcB/uhQpMY/KDgI4lImdFaLPXDrzg21NL29ZHA8+6mLXywd92C70/k1QnReNU2QLgZz1Y rFZKLuwMuhNmDyu4Fd9qYAzIIyy3mH5un8k1r7rLF7y8i8Q0/ZmIBRYykn9jum0n6pdl/y1jxTo RFFt6T+3PRZA073j5H9eG1ONxb5kzClE75Es6hSXGN2BjL+G1nEk+Q/FENzNXbRLvoY5iFqumIU gQZnuSsijadoRJD8a6w9PyrNteQSD96TtuXzzVoJLUdKgyjeYW5mKyy/z7uJWa7gXN2FdZ/wUAP i1HePB3ul8Yt3q19raTo9oeNERTXwcVWc5CHf+n4QmTuG5dgZvUKz7+MV99tv3YImqSZHsq/2z9 pgkruvU/V5NJrC04bKw== X-Authority-Analysis: v=2.4 cv=WIFPmHsR c=1 sm=1 tr=0 ts=6a72d80f cx=c_pps a=AfN7/Ok6k8XGzOShvHwTGQ==:117 a=AfN7/Ok6k8XGzOShvHwTGQ==:17 a=Sv0fKeRqtYgA:10 a=VkNPw1HP01LnGYTKEx00:22 a=RnoormkPH1_aCDwRdu11:22 a=V8glGbnc2Ofi9Qvn3v5h:22 a=VwQbUJbxAAAA:8 a=VnNF1IyMAAAA:8 a=8ZJD85kOQxw6bUDZt3oA:9 X-Proofpoint-GUID: p518RXZYhP7QoLVp4gaKCUOUg2x3O4co X-Proofpoint-ORIG-GUID: eDeq56Fk8OMNwfiGb737VVWlvKkwja-m X-Proofpoint-Spam-Info: AW1haW4tMjYwODA1MDA0NyBTYWx0ZWRfX1IpFCQCXKwec BBy1Jy/YmMXefn5MKSIRJo8UceAH/bW5AJGsBf0q8w9WjU+wzE0wEy0DlFSGg+lBdY4OAn2zrPv Vn7UQIPyqrutiackW3bN19yB7uOg26o= X-Proofpoint-Virus-Version: vendor=baseguard engine=ICAP:2.0.293,Aquarius:18.0.1176,Hydra:6.1.134,FMLib:17.12.100.49 definitions=2026-08-05_02,2026-08-04_02,2025-10-01_01 X-Proofpoint-Spam-Details: rule=outbound_notspam policy=outbound score=0 spamscore=0 lowpriorityscore=0 priorityscore=1501 phishscore=0 malwarescore=0 suspectscore=0 clxscore=1011 impostorscore=0 bulkscore=0 adultscore=0 classifier=typeunknown authscore=0 authtc= authcc= route=outbound adjust=0 reason=mlx scancount=1 engine=8.22.0-2606150000 definitions=main-2608050047 Content-Type: text/plain; charset="utf-8" The current flag-style stable write implementation uses idempotent set and clear functions. This is okay because the users for the most part just want to set or clear it once based on factors like underlying device support. However, this scheme doesn't play well when we have parallel users wanting to temporarily set and unset stable writes. For example, the upcoming RWF_WRITETHROUGH patches need stable writes to be enabled for the duration of the IO. The current scheme can lead to bugs like: RWF_WRITETHROUGH write 1 RWF_WRITETHROUGH write 2 enable stable write enable stable write submit IO disable stable write <---- WRONG submit IO disable stable write The 2nd write loses the stable write guarantee midway which is not correct. Fix this by introducing a new inflight_stable_write counter which can be used by parallel users safely. Unfortunately, due to the way the current users are designed, we cannot directly migrate them to the counter approach hence for now we will have to keep both methods till all the users adapt to the counters. Suggested-by: "Darrick J. Wong" Signed-off-by: Ojaswin Mujoo --- include/linux/fs.h | 1 + include/linux/pagemap.h | 14 +++++++++++++- 2 files changed, 14 insertions(+), 1 deletion(-) diff --git a/include/linux/fs.h b/include/linux/fs.h index 8e9bc9dda0cb..5f17aa0ed4c7 100644 --- a/include/linux/fs.h +++ b/include/linux/fs.h @@ -483,6 +483,7 @@ struct address_space { errseq_t wb_err; spinlock_t i_private_lock; struct rw_semaphore i_mmap_rwsem; + atomic_t inflight_stable_writes_count; } __attribute__((aligned(sizeof(long)))) __randomize_layout; /* * On most architectures that alignment is already the case; but diff --git a/include/linux/pagemap.h b/include/linux/pagemap.h index 2c3718d592d6..eb8b7e292478 100644 --- a/include/linux/pagemap.h +++ b/include/linux/pagemap.h @@ -306,7 +306,8 @@ static inline void mapping_clear_release_always(struct = address_space *mapping) =20 static inline bool mapping_stable_writes(const struct address_space *mappi= ng) { - return test_bit(AS_STABLE_WRITES, &mapping->flags); + return test_bit(AS_STABLE_WRITES, &mapping->flags) || + atomic_read(&mapping->inflight_stable_writes_count) > 0; } =20 static inline void mapping_set_stable_writes(struct address_space *mapping) @@ -319,6 +320,17 @@ static inline void mapping_clear_stable_writes(struct = address_space *mapping) clear_bit(AS_STABLE_WRITES, &mapping->flags); } =20 +static inline void mapping_inc_inflight_stable_writes(struct address_space= *mapping) +{ + atomic_inc(&mapping->inflight_stable_writes_count); +} + +static inline void mapping_dec_inflight_stable_writes(struct address_space= *mapping) +{ + WARN_ON_ONCE(atomic_read(&mapping->inflight_stable_writes_count) =3D=3D 0= ); + atomic_dec_if_positive(&mapping->inflight_stable_writes_count); +} + static inline void mapping_set_inaccessible(struct address_space *mapping) { /* --=20 2.55.0 From nobody Sat Oct 3 04:22:53 2026 Received: from mx0b-001b2d01.pphosted.com (mx0b-001b2d01.pphosted.com [148.163.158.5]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id B05F1330644; Wed, 5 Aug 2026 06:29:06 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=148.163.158.5 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785911348; cv=none; b=siz+PvyLsDLttJ4rwC3B8YkVu3658Ta5yF0oJj+/bxkEMfLysLal4emufc+OgDoPE8Fox82PEK0ivNnX8dMS4kNkdd1axB4YHMA+b6q5YJkgXsBZRU9MD2gYRFPRMePYwlxODa9ihfvHhd8AohggRSipkM80grfRwlkLiYImDNQ= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785911348; c=relaxed/simple; bh=UAHxJKDua1+2++RgkP4V0Qyl7MCdvywG0DGSAxQ+lqk=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=dWsJuIMqTs420ZaoUIFiWXCZEs3JSEi3Y6v2p2BxFZ6IMYg3r7dkcIHs0iWny2GkpgNyN9KMwDjEOOu20S53SGA6aRpIPHk+Cvz62LhIVsEO+ll7NhLamu3llgzPfp+siOrB+2zgbOtimYtJfbFpFCA2SwXuOWfb0Wp8NW2Wjkc= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.ibm.com; spf=pass smtp.mailfrom=linux.ibm.com; dkim=pass (2048-bit key) header.d=ibm.com header.i=@ibm.com header.b=f58NkVRW; arc=none smtp.client-ip=148.163.158.5 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.ibm.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.ibm.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=ibm.com header.i=@ibm.com header.b="f58NkVRW" Received: from pps.filterd (m0353725.ppops.net [127.0.0.1]) by mx0a-001b2d01.pphosted.com (8.18.1.11/8.18.1.11) with ESMTP id 6755mBOW2909902; Wed, 5 Aug 2026 06:28:36 GMT DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=ibm.com; h=cc :content-transfer-encoding:date:from:in-reply-to:message-id :mime-version:references:subject:to; s=pp1; bh=BstbCR8L/CfAYlmhb r9LI7ifPPjor2ZWMfhQAtVhUeU=; b=f58NkVRWqTInkQcnHf2SSWXLsuXe/L+m6 KQpuSh6Mj3ois+fZgpRLzx06CQmpV5bGl0L3midtZ7x42gq/6nNBd8jyRN1KbmTu Wj3+GHncMt5Afl6jPFlgnpsGadNNUbNwSqaIKUYm/PIzaQTmz30jjwlSpm755j9b YsZpJIojuGNQIPZVHAKmSnaoLCRjJ0I2DLQuWVBq9QfKv3PdvP1XTzRyUXdWAoaz bL1XU7hUWTlaxwHHN2a8KrsUBr+gJrL7RO136gGLPVrct71n3CFkGHqgQ7ncTnz4 kXsuDvAqviuw4dZ9+6KFBpmrzb5fKOnwZ3IQBbI91umatReza+fag== Received: from ppma12.dal12v.mail.ibm.com (dc.9e.1632.ip4.static.sl-reverse.com [50.22.158.220]) by mx0a-001b2d01.pphosted.com (PPS) with ESMTPS id 4fs77g99mc-1 (version=TLSv1.2 cipher=ECDHE-RSA-AES256-GCM-SHA384 bits=256 verify=NOT); Wed, 05 Aug 2026 06:28:35 +0000 (GMT) Received: from pps.filterd (ppma12.dal12v.mail.ibm.com [127.0.0.1]) by ppma12.dal12v.mail.ibm.com (8.18.1.7/8.18.1.7) with ESMTP id 6756QGiM028129; Wed, 5 Aug 2026 06:28:34 GMT Received: from smtprelay06.fra02v.mail.ibm.com ([9.218.2.230]) by ppma12.dal12v.mail.ibm.com (PPS) with ESMTPS id 4fsu4qnfbw-1 (version=TLSv1.2 cipher=ECDHE-RSA-AES256-GCM-SHA384 bits=256 verify=NOT); Wed, 05 Aug 2026 06:28:34 +0000 (GMT) Received: from smtpav07.fra02v.mail.ibm.com (smtpav07.fra02v.mail.ibm.com [10.20.54.106]) by smtprelay06.fra02v.mail.ibm.com (8.14.9/8.14.9/NCO v10.0) with ESMTP id 6756SXT629294948 (version=TLSv1/SSLv3 cipher=DHE-RSA-AES256-GCM-SHA384 bits=256 verify=OK); Wed, 5 Aug 2026 06:28:33 GMT Received: from smtpav07.fra02v.mail.ibm.com (unknown [127.0.0.1]) by IMSVA (Postfix) with ESMTP id DECC62004D; Wed, 5 Aug 2026 06:28:32 +0000 (GMT) Received: from smtpav07.fra02v.mail.ibm.com (unknown [127.0.0.1]) by IMSVA (Postfix) with ESMTP id 9BD5720040; Wed, 5 Aug 2026 06:28:28 +0000 (GMT) Received: from li-dc0c254c-257c-11b2-a85c-98b6c1322444.ibm.com (unknown [9.124.211.239]) by smtpav07.fra02v.mail.ibm.com (Postfix) with ESMTP; Wed, 5 Aug 2026 06:28:28 +0000 (GMT) From: Ojaswin Mujoo To: Christian Brauner , linux-fsdevel@vger.kernel.org Cc: "Darrick J . Wong" , Carlos Maiolino , Alexander Viro , Jan Kara , Matthew Wilcox , Andrew Morton , Ritesh Harjani , Zhang Yi , Christoph Hellwig , Dave Chinner , Daniel Gomez , Pankaj Raghav , Theodore Tso , linux-xfs@vger.kernel.org, linux-kernel@vger.kernel.org, linux-mm@kvack.org Subject: [RFC PATCH v3 02/11] mm: Refactor folio_clear_dirty_for_io() Date: Wed, 5 Aug 2026 11:58:08 +0530 Message-ID: <79acedfb7f3cb1814a68954974da30fd81c655bb.1785908600.git.ojaswin@linux.ibm.com> X-Mailer: git-send-email 2.55.0 In-Reply-To: References: Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-TM-AS-GCONF: 00 X-Proofpoint-Reinject: loops=2 maxloops=12 X-Proofpoint-Spam-Details-Enc: AW1haW4tMjYwODA1MDA0NyBTYWx0ZWRfX8vvbnxdDAfRP a6yid2dbDjNjKL+a4utEqt0iJZ4X6j3w7m6yICuI78XPtEaoNKgt1FSRJQ46V9ErM4VcykwAenb Qr8uoqWvzw41WmtwjQAvgJpu6dAZitDhHEuRerc8mHEv2R/iPaCBhZ4Ewz0j3QTSGCezeUwthgm QN53aN/RalRzgI/30tejS4IpokGLjU0f29YNfKCkXHrwuLtfzH+EshWE2X7/yJvplw7bmzWtavX tnx4SKeviPhRdxZbVLhg/JCZmGsSAFq+ocRYkevETa5iWiQs96VK3aYZNm135l9InyejWjl0O1c bXiuYjOA7M2vhJfhHQbw7WQ3RI6X6J0wAYFCd3F2zGD3fApwewsPO9BjrK0axBV7Zh3DxL8/qRm z3qseVb1OvJxuuEKp3cPeFkD7BcgIezQlA+ft8SDeYSQSHXqg5FEfXQhQKr4f50my6H6LUId9ic 5GHvHotRelGklvyEuxg== X-Authority-Analysis: v=2.4 cv=WIFPmHsR c=1 sm=1 tr=0 ts=6a72d814 cx=c_pps a=bLidbwmWQ0KltjZqbj+ezA==:117 a=bLidbwmWQ0KltjZqbj+ezA==:17 a=Sv0fKeRqtYgA:10 a=VkNPw1HP01LnGYTKEx00:22 a=RnoormkPH1_aCDwRdu11:22 a=V8glGbnc2Ofi9Qvn3v5h:22 a=pGLkceISAAAA:8 a=VnNF1IyMAAAA:8 a=GwQEhMSs4eZOB5FW1yQA:9 X-Proofpoint-GUID: _Cu7__fDUJK1eBU0cjcnDPFT61jAT4oL X-Proofpoint-ORIG-GUID: rCElAljz-VkEQ2Q1yGTn4bfQj2YDmix2 X-Proofpoint-Spam-Info: AW1haW4tMjYwODA1MDA0NyBTYWx0ZWRfX9P5Qk2vGZbE6 xGac7B6zFtoqopEDEkWvw/e3fN7rzlz6TgyWlX1621XLY6+VI7BXaBbzQJi1rrjhSqpQN/+9zx7 m3HoDAp6JFvJtaSlgIgFLhafxusR3gU= X-Proofpoint-Virus-Version: vendor=baseguard engine=ICAP:2.0.293,Aquarius:18.0.1176,Hydra:6.1.134,FMLib:17.12.100.49 definitions=2026-08-05_02,2026-08-04_02,2025-10-01_01 X-Proofpoint-Spam-Details: rule=outbound_notspam policy=outbound score=0 spamscore=0 lowpriorityscore=0 priorityscore=1501 phishscore=0 malwarescore=0 suspectscore=0 clxscore=1011 impostorscore=0 bulkscore=0 adultscore=0 classifier=typeunknown authscore=0 authtc= authcc= route=outbound adjust=0 reason=mlx scancount=1 engine=8.22.0-2606150000 definitions=main-2608050047 Content-Type: text/plain; charset="utf-8" Add a new __folio_clear_dirty_for_io() helper which takes an extra parameter to indicate folio_mkclean() is needed. This is in preparation of buffered writethrough support where we already do folio_mkclean() before calling into this function. Co-developed-by: Ritesh Harjani (IBM) Signed-off-by: Ritesh Harjani (IBM) Signed-off-by: Ojaswin Mujoo --- mm/page-writeback.c | 39 +++++++++++++++++++++++++-------------- 1 file changed, 25 insertions(+), 14 deletions(-) diff --git a/mm/page-writeback.c b/mm/page-writeback.c index e98748112d1e..3d184ca316a8 100644 --- a/mm/page-writeback.c +++ b/mm/page-writeback.c @@ -2856,20 +2856,12 @@ void __folio_cancel_dirty(struct folio *folio) EXPORT_SYMBOL(__folio_cancel_dirty); =20 /* - * Clear a folio's dirty flag, while caring for dirty memory accounting. - * Returns true if the folio was previously dirty. - * - * This is for preparing to put the folio under writeout. We leave - * the folio tagged as dirty in the xarray so that a concurrent - * write-for-sync can discover it via a PAGECACHE_TAG_DIRTY walk. - * The ->writepage implementation will run either folio_start_writeback() - * or folio_mark_dirty(), at which stage we bring the folio's dirty flag - * and xarray dirty tag back into sync. - * - * This incoherency between the folio's dirty flag and xarray tag is - * unfortunate, but it only exists while the folio is locked. + * Internal helper to take care of clearing dirty bit on a folio in prepar= ation + * of an IO. For some cases we might not want to do mkclean, eg, if we've + * already taken care of it, hence pass the should_mkclean flag to indicat= e if + * its needed. */ -bool folio_clear_dirty_for_io(struct folio *folio) +static bool __folio_clear_dirty_for_io(struct folio *folio, bool should_mk= clean) { struct address_space *mapping =3D folio_mapping(folio); bool ret =3D false; @@ -2906,7 +2898,7 @@ bool folio_clear_dirty_for_io(struct folio *folio) * as a serialization point for all the different * threads doing their things. */ - if (folio_mkclean(folio)) + if (should_mkclean && folio_mkclean(folio)) folio_mark_dirty(folio); /* * We carefully synchronise fault handlers against @@ -2931,6 +2923,25 @@ bool folio_clear_dirty_for_io(struct folio *folio) } return folio_test_clear_dirty(folio); } + +/* + * Clear a folio's dirty flag, while caring for dirty memory accounting. + * Returns true if the folio was previously dirty. + * + * This is for preparing to put the folio under writeout. We leave + * the folio tagged as dirty in the xarray so that a concurrent + * write-for-sync can discover it via a PAGECACHE_TAG_DIRTY walk. + * The ->writepage implementation will run either folio_start_writeback() + * or folio_mark_dirty(), at which stage we bring the folio's dirty flag + * and xarray dirty tag back into sync. + * + * This incoherency between the folio's dirty flag and xarray tag is + * unfortunate, but it only exists while the folio is locked. + */ +bool folio_clear_dirty_for_io(struct folio *folio) +{ + return __folio_clear_dirty_for_io(folio, true); +} EXPORT_SYMBOL(folio_clear_dirty_for_io); =20 static void wb_inode_writeback_start(struct bdi_writeback *wb) --=20 2.55.0 From nobody Sat Oct 3 04:22:53 2026 Received: from mx0a-001b2d01.pphosted.com (mx0a-001b2d01.pphosted.com [148.163.156.1]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 0C13A3DD503; Wed, 5 Aug 2026 06:29:18 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=148.163.156.1 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785911361; cv=none; b=I27nQDRfWAAaDi98l9JV9wG3bZdrF/XE8/+o673gZhInEvJR4KFoGc879kMT8v7JFxvFMwI69s6CnsFqJwTwl7uuwTfaT/kC1aYKtxUQDLz+Y/398yE0DCla5QcAzyNN72ZGMlEjlv13pM8xba4aWk1bFhoQsJoVLFLE8ibuYpg= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785911361; c=relaxed/simple; bh=Sb5ZNhtNCaDmDDu2xPmV2//wfB9VzDdlgQNk2NzW0zA=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=u3QRbJ7ZikmIRbSz/7in0ZUtSN4XIx9llCdqOMnJUC8j/7IrD4TyAJcFSspOUmX6bl9Kh67LJPhElZzcl1nu3GCFD6aawoV/XgUY4OvxAK0HTOWqMESk7yWfJeZrjNkFtodNgoYeK8IXpKwxuRppZObzlFob/CijP53evhz9o+Q= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.ibm.com; spf=pass smtp.mailfrom=linux.ibm.com; dkim=pass (2048-bit key) header.d=ibm.com header.i=@ibm.com header.b=s1jB/175; arc=none smtp.client-ip=148.163.156.1 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.ibm.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.ibm.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=ibm.com header.i=@ibm.com header.b="s1jB/175" Received: from pps.filterd (m0353729.ppops.net [127.0.0.1]) by mx0a-001b2d01.pphosted.com (8.18.1.11/8.18.1.11) with ESMTP id 6755lvoP1121612; Wed, 5 Aug 2026 06:28:43 GMT DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=ibm.com; h=cc :content-transfer-encoding:date:from:in-reply-to:message-id :mime-version:references:subject:to; s=pp1; bh=4gSdh3sr2ycHT3Euk mQjBtk1Q919hgqEMrec8njc2Dw=; b=s1jB/175e8VH2EsR9qoQ8dSt6XYww+ZWR eOkZJSN45KgYml+6R3cLj4t8dJqIqV0KjSl3M+36YTJS9ZqvCpHi8ouzCGdk0yrG RI+7gRuQLYwQ0W3UdAlxxjpHDLj7sNsv0FjcdrE9uP8QAjO8Kob2HTeJmLLBCIpT KmgZ5OdVJ45hzK8C69o9ZmDfH4WCqLtNygAi1nsgCPvNEwgo5cq/RT+KgAVZ+0ps vJlqRfS6259CU6M3egXpNyyxon8FTHvoW8cYGy3V8W+KcvxMPJnalWdUMl1pBhY9 0qxi1bxh6uYMFIaWSZsPiWbOHXn/y+eDaIkeSX+ty9ZNFOL77M2Vw== Received: from ppma22.wdc07v.mail.ibm.com (5c.69.3da9.ip4.static.sl-reverse.com [169.61.105.92]) by mx0a-001b2d01.pphosted.com (PPS) with ESMTPS id 4fs8fqsj8h-1 (version=TLSv1.2 cipher=ECDHE-RSA-AES256-GCM-SHA384 bits=256 verify=NOT); Wed, 05 Aug 2026 06:28:42 +0000 (GMT) Received: from pps.filterd (ppma22.wdc07v.mail.ibm.com [127.0.0.1]) by ppma22.wdc07v.mail.ibm.com (8.18.1.7/8.18.1.7) with ESMTP id 6756QHSD010639; Wed, 5 Aug 2026 06:28:41 GMT Received: from smtprelay06.fra02v.mail.ibm.com ([9.218.2.230]) by ppma22.wdc07v.mail.ibm.com (PPS) with ESMTPS id 4fsugw5ee6-1 (version=TLSv1.2 cipher=ECDHE-RSA-AES256-GCM-SHA384 bits=256 verify=NOT); Wed, 05 Aug 2026 06:28:41 +0000 (GMT) Received: from smtpav07.fra02v.mail.ibm.com (smtpav07.fra02v.mail.ibm.com [10.20.54.106]) by smtprelay06.fra02v.mail.ibm.com (8.14.9/8.14.9/NCO v10.0) with ESMTP id 6756Sds531457622 (version=TLSv1/SSLv3 cipher=DHE-RSA-AES256-GCM-SHA384 bits=256 verify=OK); Wed, 5 Aug 2026 06:28:39 GMT Received: from smtpav07.fra02v.mail.ibm.com (unknown [127.0.0.1]) by IMSVA (Postfix) with ESMTP id F07072004B; Wed, 5 Aug 2026 06:28:38 +0000 (GMT) Received: from smtpav07.fra02v.mail.ibm.com (unknown [127.0.0.1]) by IMSVA (Postfix) with ESMTP id 6D66920040; Wed, 5 Aug 2026 06:28:33 +0000 (GMT) Received: from li-dc0c254c-257c-11b2-a85c-98b6c1322444.ibm.com (unknown [9.124.211.239]) by smtpav07.fra02v.mail.ibm.com (Postfix) with ESMTP; Wed, 5 Aug 2026 06:28:33 +0000 (GMT) From: Ojaswin Mujoo To: Christian Brauner , linux-fsdevel@vger.kernel.org Cc: "Darrick J . Wong" , Carlos Maiolino , Alexander Viro , Jan Kara , Matthew Wilcox , Andrew Morton , Ritesh Harjani , Zhang Yi , Christoph Hellwig , Dave Chinner , Daniel Gomez , Pankaj Raghav , Theodore Tso , linux-xfs@vger.kernel.org, linux-kernel@vger.kernel.org, linux-mm@kvack.org Subject: [RFC PATCH v3 03/11] iomap: Add helper to revert iomap iter Date: Wed, 5 Aug 2026 11:58:09 +0530 Message-ID: X-Mailer: git-send-email 2.55.0 In-Reply-To: References: Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-TM-AS-GCONF: 00 X-Proofpoint-Reinject: loops=2 maxloops=12 X-Proofpoint-GUID: LonUA8QG8JHck5z9DjLFdbCjPxe3fRjO X-Proofpoint-ORIG-GUID: HQc6MlCSscsYcwTphpnZSkOKguor5Xs- X-Proofpoint-Spam-Info: AW1haW4tMjYwODA1MDA0NyBTYWx0ZWRfX/rVMYMlnGX9p nT8eKbRD1ttQw7DughHsM3iPyZLw/hfH+Is0CFRzdoSFJpYpBhK9emmwkrhNhl2nJsN1K759f+z Nr+R9Z3JjkQQUm1LgwKRJJnjxSMrnNU= X-Authority-Analysis: v=2.4 cv=K8cS2SWI c=1 sm=1 tr=0 ts=6a72d81b cx=c_pps a=5BHTudwdYE3Te8bg5FgnPg==:117 a=5BHTudwdYE3Te8bg5FgnPg==:17 a=Sv0fKeRqtYgA:10 a=VkNPw1HP01LnGYTKEx00:22 a=RnoormkPH1_aCDwRdu11:22 a=uAbxVGIbfxUO_5tXvNgY:22 a=VnNF1IyMAAAA:8 a=EiUwcbAk3iOqCFawNVEA:9 X-Proofpoint-Spam-Details-Enc: AW1haW4tMjYwODA1MDA0NyBTYWx0ZWRfX10VcwEXMqktr +S/UFnL70GiUKVXlYY0cCvozuvypPIJgBCHxjhDCDbIZhMm/a93vKQb28I7oPh5tS4SA8f4xH1Q DF8f4jd47DRQCLvsAxONfFpKvqjBk/voYjW02yljWfssblg5yXlCrKtP/mcAJnkRWtfRayJsKQs vMyPBk+eO9M4DQZoA4EVgThV2/9LPf3bqU9TSOZQ5j/OsTksSexIbXVo9aC80BDfyEvBkUv5AHM 1lQzlr5ZOhN7b1nV8Bc3nnD7a9Qc0p118sbmxjGk5ovAIZYphtsHxUaW8enkIDq/ljzG9C0Raul U24pScs778VG69M4qRq+ftI6voEQ/9T3dyiELKOJUodJvvteISS/I1N9o3gVod+7Daft6QkBl12 mkP99TM42MUKhQps3TZpLLNlhbMRgq7yZpq0cijN2T7dt33iV7Y61bHCmJLAlEhb82r80+3JqQq GqPcb1SB0QR0mNpqJLA== X-Proofpoint-Virus-Version: vendor=baseguard engine=ICAP:2.0.293,Aquarius:18.0.1176,Hydra:6.1.134,FMLib:17.12.100.49 definitions=2026-08-05_02,2026-08-04_02,2025-10-01_01 X-Proofpoint-Spam-Details: rule=outbound_notspam policy=outbound score=0 clxscore=1015 spamscore=0 impostorscore=0 bulkscore=0 priorityscore=1501 lowpriorityscore=0 malwarescore=0 phishscore=0 suspectscore=0 adultscore=0 classifier=typeunknown authscore=0 authtc= authcc= route=outbound adjust=0 reason=mlx scancount=1 engine=8.22.0-2606150000 definitions=main-2608050047 Content-Type: text/plain; charset="utf-8" Add helper to revert back the iomap iter by count bytes. This will be used by upcoming RWF_WRITETHROUGH feature to revert the state of the iomap in case an error is encountered. Writethrough can process and park folios so that they can be batched together for IO later, this also advances the iomap iter. However if we encounter any error, we would want to cancel the IO on those processed folios. In this case, we want to revert the iter back so the upper layers can know how much we were actually able to write before failure. Signed-off-by: Ojaswin Mujoo --- fs/iomap/iter.c | 9 +++++++++ include/linux/iomap.h | 1 + 2 files changed, 10 insertions(+) diff --git a/fs/iomap/iter.c b/fs/iomap/iter.c index c445a38b6285..8ea849e7ccef 100644 --- a/fs/iomap/iter.c +++ b/fs/iomap/iter.c @@ -32,6 +32,15 @@ int iomap_iter_advance(struct iomap_iter *iter, u64 coun= t) return 0; } =20 +int iomap_iter_revert(struct iomap_iter *iter, u64 count) +{ + if (WARN_ON_ONCE(count > iter->pos)) + return -EIO; + iter->pos -=3D count; + iter->len +=3D count; + return 0; +} + static inline void iomap_iter_done(struct iomap_iter *iter) { WARN_ON_ONCE(iter->iomap.offset > iter->pos); diff --git a/include/linux/iomap.h b/include/linux/iomap.h index 8c754eb974fb..1c8f13948f81 100644 --- a/include/linux/iomap.h +++ b/include/linux/iomap.h @@ -275,6 +275,7 @@ struct iomap_iter { =20 int iomap_iter(struct iomap_iter *iter, const struct iomap_ops *ops); int iomap_iter_advance(struct iomap_iter *iter, u64 count); +int iomap_iter_revert(struct iomap_iter *iter, u64 count); =20 /** * iomap_length_trim - trimmed length of the current iomap iteration --=20 2.55.0 From nobody Sat Oct 3 04:22:53 2026 Received: from mx0a-001b2d01.pphosted.com (mx0a-001b2d01.pphosted.com [148.163.156.1]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 3EC7E3DCD80; Wed, 5 Aug 2026 06:29:17 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=148.163.156.1 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785911360; cv=none; b=qZVf7XgZs8PQ/0JpH1MIwiGGYDgqQ5vvxG174szqMgqc6lKwhbN5L970LM4gVYxCIkGOrmkX6OXtEx/fs9FvWKjUryIvbjOFOWpfmioMNJyO4dsfk1mDrQv3IG67byLVlSr+LIVccS7bLFOJvSOwD5YvHStJnqHivh3rqQ1F/VI= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785911360; c=relaxed/simple; bh=Ioh6XhoBQbHS2rflLBrH+uePfanYFNrptMQr/+txJB4=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=SwA2ZMcyaHtUD47FlNLp0k5f//xcls6N7Kqvkf9denJN+C82pcPDMVRumMu5NeGDlajOYZUwy7K+IaPpsxnRZzryvGR+sCfFYk+XaQ6L1IlpQdnqypWdMOVrZOLNXzOeHZsHLEvzpA8gHS+F4Dn3TYt83sR8t4Ce8Lt2e1t+ydY= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.ibm.com; spf=pass smtp.mailfrom=linux.ibm.com; dkim=pass (2048-bit key) header.d=ibm.com header.i=@ibm.com header.b=sUdiUIGO; arc=none smtp.client-ip=148.163.156.1 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.ibm.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.ibm.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=ibm.com header.i=@ibm.com header.b="sUdiUIGO" Received: from pps.filterd (m0360083.ppops.net [127.0.0.1]) by mx0a-001b2d01.pphosted.com (8.18.1.11/8.18.1.11) with ESMTP id 6755llhl2973321; Wed, 5 Aug 2026 06:28:48 GMT DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=ibm.com; h=cc :content-transfer-encoding:date:from:in-reply-to:message-id :mime-version:references:subject:to; s=pp1; bh=voH7uIva5aChHIbaJ 51UAuhU5U/XSg80MnqMWEh84Cw=; b=sUdiUIGOB3EHGmJa2AB9OOtTlrazR7qX5 MO3EfH9iR9WEHfk1qFZV1q8F9YddEwzPsiEI5eZ4qEEzNoyH/4EoEupuOOCwBL3b AJRaQJ8bbefz9JbExrNn0Upo8yQCFll6RMlfg2SwxmqPAqdHfEAltT/LDPUppcnC YjyJBAkqsVTPxFzBig9R9A1Wa+KrCp0XD3qy5x5o4KXFgdey3jCQCnFQIzsvqQpI Zn9r3K7eHR9XwK7As66NVJYJN80vDm7O21+3i/fhESR3OJdLaON69mpn0RHN6LtH qLgFzB3hPPCT0x1vJ/Ewjhm87qCq8LZN8Mh19XOxbvVUVv8pHhOsA== Received: from ppma23.wdc07v.mail.ibm.com (5d.69.3da9.ip4.static.sl-reverse.com [169.61.105.93]) by mx0a-001b2d01.pphosted.com (PPS) with ESMTPS id 4fs8a41mr5-1 (version=TLSv1.2 cipher=ECDHE-RSA-AES256-GCM-SHA384 bits=256 verify=NOT); Wed, 05 Aug 2026 06:28:47 +0000 (GMT) Received: from pps.filterd (ppma23.wdc07v.mail.ibm.com [127.0.0.1]) by ppma23.wdc07v.mail.ibm.com (8.18.1.7/8.18.1.7) with ESMTP id 6756QIBR019933; Wed, 5 Aug 2026 06:28:46 GMT Received: from smtprelay04.fra02v.mail.ibm.com ([9.218.2.228]) by ppma23.wdc07v.mail.ibm.com (PPS) with ESMTPS id 4fsvmhd8rt-1 (version=TLSv1.2 cipher=ECDHE-RSA-AES256-GCM-SHA384 bits=256 verify=NOT); Wed, 05 Aug 2026 06:28:46 +0000 (GMT) Received: from smtpav07.fra02v.mail.ibm.com (smtpav07.fra02v.mail.ibm.com [10.20.54.106]) by smtprelay04.fra02v.mail.ibm.com (8.14.9/8.14.9/NCO v10.0) with ESMTP id 6756Sis116712168 (version=TLSv1/SSLv3 cipher=DHE-RSA-AES256-GCM-SHA384 bits=256 verify=OK); Wed, 5 Aug 2026 06:28:44 GMT Received: from smtpav07.fra02v.mail.ibm.com (unknown [127.0.0.1]) by IMSVA (Postfix) with ESMTP id 49E452004B; Wed, 5 Aug 2026 06:28:44 +0000 (GMT) Received: from smtpav07.fra02v.mail.ibm.com (unknown [127.0.0.1]) by IMSVA (Postfix) with ESMTP id 9098420040; Wed, 5 Aug 2026 06:28:39 +0000 (GMT) Received: from li-dc0c254c-257c-11b2-a85c-98b6c1322444.ibm.com (unknown [9.124.211.239]) by smtpav07.fra02v.mail.ibm.com (Postfix) with ESMTP; Wed, 5 Aug 2026 06:28:39 +0000 (GMT) From: Ojaswin Mujoo To: Christian Brauner , linux-fsdevel@vger.kernel.org Cc: "Darrick J . Wong" , Carlos Maiolino , Alexander Viro , Jan Kara , Matthew Wilcox , Andrew Morton , Ritesh Harjani , Zhang Yi , Christoph Hellwig , Dave Chinner , Daniel Gomez , Pankaj Raghav , Theodore Tso , linux-xfs@vger.kernel.org, linux-kernel@vger.kernel.org, linux-mm@kvack.org, Dave Chinner Subject: [RFC PATCH v3 04/11] iomap: Add initial support for buffered RWF_WRITETHROUGH Date: Wed, 5 Aug 2026 11:58:10 +0530 Message-ID: X-Mailer: git-send-email 2.55.0 In-Reply-To: References: Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-TM-AS-GCONF: 00 X-Proofpoint-Reinject: loops=2 maxloops=12 X-Authority-Analysis: v=2.4 cv=E6P9Y6dl c=1 sm=1 tr=0 ts=6a72d820 cx=c_pps a=3Bg1Hr4SwmMryq2xdFQyZA==:117 a=3Bg1Hr4SwmMryq2xdFQyZA==:17 a=Sv0fKeRqtYgA:10 a=VkNPw1HP01LnGYTKEx00:22 a=RnoormkPH1_aCDwRdu11:22 a=iQ6ETzBq9ecOQQE5vZCe:22 a=VwQbUJbxAAAA:8 a=pGLkceISAAAA:8 a=VnNF1IyMAAAA:8 a=JjhpDi1jizst4QKIiw0A:9 X-Proofpoint-ORIG-GUID: mwghkPdwfzsJfGcGOZ_fqG1ncJdTv1Pj X-Proofpoint-GUID: LUHfteGxd35IVFMRofxcJa0EEmXMboVH X-Proofpoint-Spam-Info: AW1haW4tMjYwODA1MDA0NyBTYWx0ZWRfXxfKzAn4SYSqi AHCFW6Ry5Cs5j1qW4hliuc302G89jTpip8bUHJYtxhmVKv81ygdUha8p/njrEAvTi2KPCwfI1Pk DdpKy0q+l/JmkUbAkbgdpePThtSIQDg= X-Proofpoint-Spam-Details-Enc: AW1haW4tMjYwODA1MDA0NyBTYWx0ZWRfX38OF/grlxWSQ DugseEhojoEi16L6MIlXz5ltalixQ7qYJhVWwurL0+6BmLXCQ0KjcReXhaJtgovtcyMPmZ1JGrf t/z2RtUahLAiW2xMsBfANkLkKwM4tVDOKGbcI8kMexEwuYZY+cf3PtZbB0zhViVjSSe6TdZP7jS /g7BnO3VVt0xIYP4S620nNxPNeAK5+LH9bt02kP5Ot0FAlmYtCgSzZkm4QzGWb3i0xtvAACO2GK M5isEba1LeFVUv8ixEyoicozLLJrGPdy9LkhvV7JC/o7Jw7twqJN3PjHRudOC6n0ylBcoJ6AlvY LAceMtE3QBY5Fz0ZEZcDFRtwPJd8apQnkZKui8iLTUfvNVP506QV3QuJzGptf/ZYES7Nh2AICDk sSTCmR3q7rIB3vHyjhQ/x6Br5bG1FlMSN8PtJVlj9x0l2elmGj7nZis+J+pIAdAsN5WsXjqfDF1 AZndSyo6uYKeXAWupWg== X-Proofpoint-Virus-Version: vendor=baseguard engine=ICAP:2.0.293,Aquarius:18.0.1176,Hydra:6.1.134,FMLib:17.12.100.49 definitions=2026-08-05_02,2026-08-04_02,2025-10-01_01 X-Proofpoint-Spam-Details: rule=outbound_notspam policy=outbound score=0 bulkscore=0 clxscore=1011 lowpriorityscore=0 priorityscore=1501 suspectscore=0 adultscore=0 spamscore=0 malwarescore=0 impostorscore=0 phishscore=0 classifier=typeunknown authscore=0 authtc= authcc= route=outbound adjust=0 reason=mlx scancount=1 engine=8.22.0-2606150000 definitions=main-2608050047 Content-Type: text/plain; charset="utf-8" This adds initial support for performing buffered non-aio RWF_WRITETHROUGH write. The rough flow for a writethrough write is as follows: 1. Acquire inode lock 2. initialize writethrough context (wt_ctx) and mark mapping as stable. 3. Start the iomap_iter() loop. For each iomap: 3.1. Acquire folio and folio_lock. 3.2. perform memcpy from user buffer to the folio and mark it dirty 3.3. Wait for any current writeback to complete and then call folio_mkclean() to prevent mmap writes from changing it. 3.4. Start writeback on the folio 3.5. Add the folio range under write to wt_ctx->bvec and folio_unlock() 3.6. If bvec is full, submit the current bvecs for IO. 3.7. Repeat 3.2 to 3.6 till the whole iomap is processed. Submit the final set of bvecs for IO. 4. Repeat step 3 till we have no more data to write. 5. Finally, sleep in the syscall thread till all the IOs are completed (refcount =3D=3D 0). Once that happens, the end io handler will wake us up. 6. Upon waking up, call fs ->end_io() callback (which updates inode size), record any errors and return. 7. inode_unlock() This design gives buffered writethrough the same semantics as dio and any error in the IO is directly returned to the caller. However, the users should note that an error should be treated equivalent of a buffered IO fsync error, since we can't always guarantee that the page cache is consistent. The design has deliberately open coded the IO submission and completion flow (inspired by dio) rather than reusing the dio functions as accommodating buffered writethrough logic in dio code was polluting it with too many if else conditionals and special cases. Suggested-by: Jan Kara Suggested-by: Dave Chinner Co-developed-by: Ritesh Harjani (IBM) Signed-off-by: Ritesh Harjani (IBM) Signed-off-by: Ojaswin Mujoo --- fs/iomap/buffered-io.c | 437 ++++++++++++++++++++++++++++++++++++++++ include/linux/fs.h | 11 + include/linux/iomap.h | 42 ++++ include/linux/pagemap.h | 1 + include/uapi/linux/fs.h | 5 +- mm/page-writeback.c | 10 + 6 files changed, 505 insertions(+), 1 deletion(-) diff --git a/fs/iomap/buffered-io.c b/fs/iomap/buffered-io.c index 0a5ebfda90f1..3178e8c0fa13 100644 --- a/fs/iomap/buffered-io.c +++ b/fs/iomap/buffered-io.c @@ -10,6 +10,7 @@ #include #include #include +#include #include "internal.h" #include "trace.h" =20 @@ -1185,6 +1186,360 @@ static bool iomap_write_end(struct iomap_iter *iter= , size_t len, size_t copied, return __iomap_write_end(iter->inode, pos, len, copied, folio); } =20 +static ssize_t iomap_writethrough_complete(struct iomap_writethrough_ctx *= wt_ctx) +{ + struct kiocb *iocb =3D wt_ctx->iocb; + struct inode *inode =3D wt_ctx->inode; + ssize_t ret =3D wt_ctx->error; + + if (wt_ctx->end_io) { + int err =3D wt_ctx->end_io(wt_ctx, wt_ctx->written, + wt_ctx->error, + wt_ctx->flags); + if (err) + ret =3D err; + } + + mapping_dec_inflight_stable_writes(inode->i_mapping); + + if (!ret) { + ret =3D wt_ctx->written; + iocb->ki_pos +=3D ret; + } + + kfree(wt_ctx); + return ret; +} + +static void iomap_writethrough_done(struct iomap_writethrough_ctx *wt_ctx) +{ + struct task_struct *waiter =3D wt_ctx->waiter; + + WRITE_ONCE(wt_ctx->waiter, NULL); + blk_wake_io_task(waiter); +} + +static void iomap_writethrough_bio_end_io(struct bio *bio) +{ + struct iomap_writethrough_ctx *wt_ctx =3D bio->bi_private; + struct folio_iter fi; + + if (bio->bi_status) + cmpxchg(&wt_ctx->error, 0, + blk_status_to_errno(bio->bi_status)); + bio_for_each_folio_all(fi, bio) + folio_end_writeback(fi.folio); + + bio_put(bio); + if (atomic_dec_and_test(&wt_ctx->ref)) + iomap_writethrough_done(wt_ctx); +} + +static int +iomap_writethrough_submit_bio(struct iomap_writethrough_ctx *wt_ctx, + struct iomap *iomap, + const struct iomap_writethrough_ops *wt_ops, int error) +{ + struct bio *bio; + unsigned int i; + u64 len =3D 0; + blk_opf_t opf =3D REQ_OP_WRITE | REQ_SYNC | REQ_IDLE; + + if (!wt_ctx->nr_bvecs) + goto exit; + + for (i =3D 0; i < wt_ctx->nr_bvecs; i++) + len +=3D wt_ctx->bvec[i].bv_len; + + bio =3D bio_alloc(iomap->bdev, wt_ctx->nr_bvecs, opf, GFP_NOFS); + bio->bi_iter.bi_sector =3D iomap_sector(iomap, wt_ctx->bio_pos); + bio->bi_end_io =3D iomap_writethrough_bio_end_io; + bio->bi_private =3D wt_ctx; + + for (i =3D 0; i < wt_ctx->nr_bvecs; i++) + __bio_add_page(bio, wt_ctx->bvec[i].bv_page, + wt_ctx->bvec[i].bv_len, + wt_ctx->bvec[i].bv_offset); + + if (!error && wt_ops->writethrough_submit) + error =3D wt_ops->writethrough_submit(wt_ctx->inode, iomap, + wt_ctx->bio_pos, len); + + + atomic_inc(&wt_ctx->ref); + + /* + * In case of error we still need the I/O completion to run so we can + * release references and end writeback on the folios. + */ + if (error) { + bio->bi_status =3D errno_to_blk_status(error); + bio_endio(bio); + return error; + } + + submit_bio(bio); + wt_ctx->nr_bvecs =3D 0; + +exit: + return 0; +} + +/* + * Submit any pending bvecs as a bio and account the written bytes. + * On failure, iomap_writethrough_submit_bio() has already called the + * endio completion handler to record the error. + */ +static int +iomap_writethrough_try_submit(struct iomap_writethrough_ctx *wt_ctx, + struct iomap *iomap, + const struct iomap_writethrough_ops *wt_ops, + ssize_t *pending) +{ + int ret =3D iomap_writethrough_submit_bio(wt_ctx, iomap, wt_ops, 0); + + if (ret < 0) + return ret; + wt_ctx->written +=3D *pending; + *pending =3D 0; + return 0; +} + +/** + * iomap_writethrough_begin - prepare the various structures for writethro= ugh + * @folio: folio to prepare for writethrough + * @off: offset of write within folio + * @len: len of write within folio + * + * This function does the major preparation work needed before starting the + * writethrough. The main task is to prepare folio for writeththrough by b= locking + * mmap writes and setting writeback on it. Further, we must clear the wri= te range + * to non-dirty. If this results in the complete folio becoming non-dirty,= then we + * need to clear the master dirty bit. + */ +static void iomap_folio_prepare_writethrough(struct folio *folio, size_t o= ff, + size_t len) +{ + bool fully_written; + u64 zero =3D 0; + + if (folio_test_writeback(folio)) + folio_wait_writeback(folio); + + if (folio_mkclean(folio)) + folio_mark_dirty(folio); + + /* + * We might either write through the complete folio or a partial folio + * writethrough might result in all blocks becoming non-dirty, so we need= to + * check and mark the folio clean if that is the case. + */ + fully_written =3D (off =3D=3D 0 && len =3D=3D folio_size(folio)); + iomap_clear_range_dirty(folio, off, len); + if (fully_written || + !iomap_find_dirty_range(folio, &zero, folio_size(folio))) + folio_clear_dirty_for_writethrough(folio); + + folio_start_writeback(folio); +} + +/** + * iomap_writethrough_iter - perform RWF_WRITETHROUGH buffered write + * @wt_ctx: writethrough context + * @iter: iomap iter holding mapping information + * @i: iov_iter for write + * @wt_ops: the fs callbacks needed for writethrough + * + * This function copies the user buffer to folio similar to usual buffered + * IO path, with the difference that we immediately issue the IO. For this= we + * utilize IO submission and completion mechanism that is inspired by dio. + * + * Folio handling note: We might be writing through a partial folio so we = need + * to be careful to not clear the folio dirty bit unless there are no dirt= y blocks + * in the folio after the writethrough. + */ +static int iomap_writethrough_iter(struct iomap_writethrough_ctx *wt_ctx, + struct iomap_iter *iter, struct iov_iter *i, + const struct iomap_writethrough_ops *wt_ops) + +{ + ssize_t total_written =3D 0, pending =3D 0; + loff_t submit_start_pos; + int status =3D 0; + struct address_space *mapping =3D iter->inode->i_mapping; + size_t chunk =3D mapping_max_folio_size(mapping); + unsigned int bdp_flags =3D (iter->flags & IOMAP_NOWAIT) ? BDP_ASYNC : 0; + unsigned int bs =3D i_blocksize(iter->inode); + + /* copied over based on how DIO handles these flags */ + if (iter->iomap.type =3D=3D IOMAP_UNWRITTEN) + wt_ctx->flags |=3D IOMAP_DIO_UNWRITTEN; + if (iter->iomap.flags & IOMAP_F_SHARED) + wt_ctx->flags |=3D IOMAP_DIO_COW; + + if (!(iter->flags & IOMAP_WRITETHROUGH)) + return -EINVAL; + + /* + * IOMAP_INLINE mappings have NULL bdev and would cause + * iomap_sector() to dereference invalid memory. Reject them. + */ + if (iter->iomap.type =3D=3D IOMAP_INLINE) + return -EINVAL; + + do { + struct folio *folio; + size_t offset; /* Offset into folio */ + loff_t old_size; + u64 bytes; /* Bytes to write to folio */ + size_t copied; /* Bytes copied from user */ + u64 written; /* Bytes have been written */ + loff_t pos; + size_t off_aligned, len_aligned; + + bytes =3D iov_iter_count(i); +retry: + offset =3D iter->pos & (chunk - 1); + bytes =3D min(chunk - offset, bytes); + status =3D balance_dirty_pages_ratelimited_flags(mapping, + bdp_flags); + if (unlikely(status)) + break; + + /* + * If completions already occurred and reported errors, give up + * now and don't bother submitting more bios. + */ + status =3D data_race(wt_ctx->error); + if (unlikely(status)) { + wt_ctx->nr_bvecs =3D 0; + break; + } + + if (bytes > iomap_length(iter)) + bytes =3D iomap_length(iter); + + /* + * Bring in the user page that we'll copy from _first_. + * Otherwise there's a nasty deadlock on copying from the + * same page as we're writing to, without it being marked + * up-to-date. + * + * For async buffered writes the assumption is that the user + * page has already been faulted in. This can be optimized by + * faulting the user page. + */ + if (unlikely(fault_in_iov_iter_readable(i, bytes) =3D=3D bytes)) { + status =3D -EFAULT; + break; + } + + status =3D iomap_write_begin(iter, wt_ops->write_ops, &folio, + &offset, &bytes); + if (unlikely(status)) { + iomap_write_failed(iter->inode, iter->pos, bytes); + break; + } + if (iter->iomap.flags & IOMAP_F_STALE) + break; + + pos =3D iter->pos; + + if (mapping_writably_mapped(mapping)) + flush_dcache_folio(folio); + + copied =3D copy_folio_from_iter_atomic(folio, offset, bytes, i); + written =3D iomap_write_end(iter, bytes, copied, folio) ? + copied : 0; + + old_size =3D iter->inode->i_size; + if (pos + written > old_size) + i_size_write(iter->inode, pos + written); + + if (!written) + goto put_folio; + + off_aligned =3D round_down(offset, bs); + len_aligned =3D round_up(offset + written, bs) - off_aligned; + + iomap_folio_prepare_writethrough(folio, off_aligned, + len_aligned); + + if (!wt_ctx->nr_bvecs) { + wt_ctx->bio_pos =3D round_down(pos, bs); + submit_start_pos =3D pos; + } + + bvec_set_folio(&wt_ctx->bvec[wt_ctx->nr_bvecs], folio, + len_aligned, off_aligned); + wt_ctx->nr_bvecs++; + +put_folio: + __iomap_put_folio(iter, wt_ops->write_ops, written, folio); + + if (old_size < pos) + pagecache_isize_extended(iter->inode, old_size, pos); + + cond_resched(); + if (unlikely(written =3D=3D 0)) { + iomap_write_failed(iter->inode, pos, bytes); + iov_iter_revert(i, copied); + + if (chunk > PAGE_SIZE) + chunk /=3D 2; + if (copied) { + bytes =3D copied; + goto retry; + } + } else { + total_written +=3D written; + pending +=3D written; + iomap_iter_advance(iter, written); + } + + /* + * If we fail to submit the bio, we immediately call the + * IO completion handler that records the error. We + * shall not retry anymore cause this could lead to + * infinite loops in case of non-transient errors. + */ + if (wt_ctx->nr_bvecs =3D=3D wt_ctx->max_bvecs) { + status =3D iomap_writethrough_try_submit(wt_ctx, + &iter->iomap, wt_ops, &pending); + if (status) + goto submit_failed; + } + + } while (iov_iter_count(i) && iomap_length(iter)); + + if (wt_ctx->nr_bvecs) { + status =3D iomap_writethrough_try_submit(wt_ctx, + &iter->iomap, wt_ops, &pending); + if (status) + goto submit_failed; + } + + /* + * In case of an error, we only consider the bytes we were actually able + * to submit IO for as valid data and revert the iters accordingly + */ + if (status) { + /* + * we still need to run the endio completion for cleanup work + * hence call the below helper to take care of it, if we haven't + * already done so. We can ignore the return value here. + */ + iomap_writethrough_submit_bio(wt_ctx, &iter->iomap, wt_ops, status); + +submit_failed: + iomap_write_failed(iter->inode, submit_start_pos, pending); + iomap_iter_revert(iter, pending); + iov_iter_revert(i, pending); + } + + return status; +} + static int iomap_write_iter(struct iomap_iter *iter, struct iov_iter *i, const struct iomap_write_ops *write_ops) { @@ -1345,6 +1700,88 @@ int iomap_fsverity_write(struct file *file, loff_t p= os, size_t length, } EXPORT_SYMBOL_GPL(iomap_fsverity_write); =20 +ssize_t iomap_file_writethrough_write(struct kiocb *iocb, struct iov_iter = *i, + const struct iomap_writethrough_ops *wt_ops, + void *private) +{ + struct inode *inode =3D iocb->ki_filp->f_mapping->host; + struct iomap_iter iter =3D { + .inode =3D inode, + .pos =3D iocb->ki_pos, + .len =3D iov_iter_count(i), + .flags =3D IOMAP_WRITE | IOMAP_WRITETHROUGH, + .private =3D private, + }; + struct iomap_writethrough_ctx *wt_ctx; + unsigned int max_bvecs; + ssize_t ret; + struct blk_plug plug; + size_t min_folio_bytes =3D PAGE_SIZE + << mapping_min_folio_order(inode->i_mapping); + + /* + * For now we don't support any other flag with WRITETHROUGH + */ + if (!(iocb->ki_flags & IOCB_WRITETHROUGH)) + return -EINVAL; + if (iocb->ki_flags & (IOCB_DONTCACHE)) + return -EINVAL; + if (iocb_is_dsync(iocb)) + /* D_SYNC support not implemented yet */ + return -EOPNOTSUPP; + if (!is_sync_kiocb(iocb)) + /* aio support not implemented yet */ + return -EOPNOTSUPP; + + /* + * +1 to max bvecs to account for unaligned write spanning multiple + * folios. Guard against overflow since iov_iter_count() returns size_t. + */ + max_bvecs =3D (unsigned int)min_t( + size_t, DIV_ROUND_UP(iov_iter_count(i), min_folio_bytes) + 1, + BIO_MAX_VECS); + + wt_ctx =3D kzalloc(struct_size(wt_ctx, bvec, max_bvecs), GFP_NOFS); + if (!wt_ctx) + return -ENOMEM; + + wt_ctx->iocb =3D iocb; + wt_ctx->inode =3D inode; + wt_ctx->end_io =3D wt_ops->end_io; + wt_ctx->old_i_size =3D i_size_read(inode); + wt_ctx->max_bvecs =3D max_bvecs; + atomic_set(&wt_ctx->ref, 1); + wt_ctx->waiter =3D current; + + mapping_inc_inflight_stable_writes(inode->i_mapping); + + blk_start_plug(&plug); + + while ((ret =3D iomap_iter(&iter, wt_ops->ops)) > 0) { + WARN_ON(iter.iomap.type !=3D IOMAP_UNWRITTEN && + iter.iomap.type !=3D IOMAP_MAPPED); + iter.status =3D iomap_writethrough_iter(wt_ctx, &iter, i, wt_ops); + } + + blk_finish_plug(&plug); + + if (ret < 0) + cmpxchg(&wt_ctx->error, 0, ret); + + if (!atomic_dec_and_test(&wt_ctx->ref)) { + for (;;) { + set_current_state(TASK_UNINTERRUPTIBLE); + if (!READ_ONCE(wt_ctx->waiter)) + break; + blk_io_schedule(); + } + __set_current_state(TASK_RUNNING); + } + + return iomap_writethrough_complete(wt_ctx); +} +EXPORT_SYMBOL_GPL(iomap_file_writethrough_write); + static void iomap_write_delalloc_ifs_punch(struct inode *inode, struct folio *folio, loff_t start_byte, loff_t end_byte, struct iomap *iomap, iomap_punch_t punch) diff --git a/include/linux/fs.h b/include/linux/fs.h index 5f17aa0ed4c7..bff09a5f90e8 100644 --- a/include/linux/fs.h +++ b/include/linux/fs.h @@ -344,6 +344,7 @@ struct readahead_control; #define IOCB_ATOMIC (__force int) RWF_ATOMIC #define IOCB_DONTCACHE (__force int) RWF_DONTCACHE #define IOCB_NOSIGNAL (__force int) RWF_NOSIGNAL +#define IOCB_WRITETHROUGH (__force int) RWF_WRITETHROUGH =20 /* non-RWF related bits - start at 16 */ #define IOCB_EVENTFD (1 << 16) @@ -1980,6 +1981,8 @@ struct file_operations { #define FOP_ASYNC_LOCK ((__force fop_flags_t)(1 << 6)) /* File system supports uncached read/write buffered IO */ #define FOP_DONTCACHE ((__force fop_flags_t)(1 << 7)) +/* File system supports write through buffered IO */ +#define FOP_WRITETHROUGH ((__force fop_flags_t)(1 << 8)) =20 /* Wrap a directory iterator that needs exclusive inode access */ int wrap_directory_iterator(struct file *, struct dir_context *, @@ -3464,6 +3467,14 @@ static inline int kiocb_set_rw_flags(struct kiocb *k= i, rwf_t flags, if (IS_DAX(ki->ki_filp->f_mapping->host)) return -EOPNOTSUPP; } + if (flags & RWF_WRITETHROUGH) { + /* file system must support it */ + if (!(ki->ki_filp->f_op->fop_flags & FOP_WRITETHROUGH)) + return -EOPNOTSUPP; + /* DAX mappings not supported */ + if (IS_DAX(ki->ki_filp->f_mapping->host)) + return -EOPNOTSUPP; + } kiocb_flags |=3D (__force int) (flags & RWF_SUPPORTED); if (flags & RWF_SYNC) kiocb_flags |=3D IOCB_DSYNC; diff --git a/include/linux/iomap.h b/include/linux/iomap.h index 1c8f13948f81..427a2763221c 100644 --- a/include/linux/iomap.h +++ b/include/linux/iomap.h @@ -212,6 +212,7 @@ struct iomap_write_ops { #endif /* CONFIG_FS_DAX */ #define IOMAP_ATOMIC (1 << 9) /* torn-write protection */ #define IOMAP_DONTCACHE (1 << 10) +#define IOMAP_WRITETHROUGH (1 << 11) =20 /* * Return the existing mapping at pos, or reserve space starting at pos fo= r up @@ -561,6 +562,29 @@ struct iomap_writepage_ctx { void *wb_ctx; /* pending writeback context */ }; =20 +struct iomap_writethrough_ctx { + struct kiocb *iocb; + struct inode *inode; + loff_t old_i_size; + loff_t new_i_size; + loff_t pos; + size_t written; + atomic_t ref; + unsigned int flags; + int error; + + /* used during submission and for non-aio completion */ + struct task_struct *waiter; + int (*end_io)(struct iomap_writethrough_ctx *wt_ctx, ssize_t size, + int error, unsigned int flags); + + loff_t bio_pos; + unsigned int nr_bvecs; + unsigned int max_bvecs; + struct bio_vec bvec[]; + +}; + struct iomap_ioend *iomap_init_ioend(struct inode *inode, struct bio *bio, loff_t file_offset, u16 ioend_flags); struct iomap_ioend *iomap_split_ioend(struct iomap_ioend *ioend, @@ -751,6 +775,24 @@ static __always_inline ssize_t iomap_dio_read_simple(s= truct kiocb *iocb, return __iomap_dio_read_simple(iocb, iter, &iomi); } =20 +/* + * In writethrough, we copy user data to folio first and then send the fol= io + * to writeback via dio path. To achieve this, we need callbacks from ioma= p_ops + * and iomap_write_ops. This struct packs them together. + */ +struct iomap_writethrough_ops { + const struct iomap_ops *ops; + const struct iomap_write_ops *write_ops; + int (*writethrough_submit)(struct inode *inode, struct iomap *iomap, + loff_t offset, u64 len); + int (*end_io)(struct iomap_writethrough_ctx *iocb, ssize_t size, + int error, unsigned int flags); +}; + +ssize_t iomap_file_writethrough_write(struct kiocb *iocb, struct iov_iter = *i, + const struct iomap_writethrough_ops *wt_ops, + void *private); + #ifdef CONFIG_SWAP struct file; struct swap_info_struct; diff --git a/include/linux/pagemap.h b/include/linux/pagemap.h index eb8b7e292478..b20e38cc0fa0 100644 --- a/include/linux/pagemap.h +++ b/include/linux/pagemap.h @@ -1270,6 +1270,7 @@ static inline void folio_cancel_dirty(struct folio *f= olio) __folio_cancel_dirty(folio); } bool folio_clear_dirty_for_io(struct folio *folio); +bool folio_clear_dirty_for_writethrough(struct folio *folio); bool clear_page_dirty_for_io(struct page *page); void folio_invalidate(struct folio *folio, size_t offset, size_t length); bool noop_dirty_folio(struct address_space *mapping, struct folio *folio); diff --git a/include/uapi/linux/fs.h b/include/uapi/linux/fs.h index bd87262f2e34..9c8d91b926a7 100644 --- a/include/uapi/linux/fs.h +++ b/include/uapi/linux/fs.h @@ -451,10 +451,13 @@ typedef int __bitwise __kernel_rwf_t; /* prevent pipe and socket writes from raising SIGPIPE */ #define RWF_NOSIGNAL ((__force __kernel_rwf_t)0x00000100) =20 +/* buffered IO that is asynchronously written through to disk after write = */ +#define RWF_WRITETHROUGH ((__force __kernel_rwf_t)0x00000200) + /* mask of flags supported by the kernel */ #define RWF_SUPPORTED (RWF_HIPRI | RWF_DSYNC | RWF_SYNC | RWF_NOWAIT |\ RWF_APPEND | RWF_NOAPPEND | RWF_ATOMIC |\ - RWF_DONTCACHE | RWF_NOSIGNAL) + RWF_DONTCACHE | RWF_NOSIGNAL | RWF_WRITETHROUGH) =20 #define PROCFS_IOCTL_MAGIC 'f' =20 diff --git a/mm/page-writeback.c b/mm/page-writeback.c index 3d184ca316a8..d4f2bb60cb47 100644 --- a/mm/page-writeback.c +++ b/mm/page-writeback.c @@ -2944,6 +2944,16 @@ bool folio_clear_dirty_for_io(struct folio *folio) } EXPORT_SYMBOL(folio_clear_dirty_for_io); =20 +/* + * Clear folio dirty in preparation of writethrough. Note that for writeth= rough + * we have already done folkio_mkclean so we avoid it here + */ +bool folio_clear_dirty_for_writethrough(struct folio *folio) +{ + return __folio_clear_dirty_for_io(folio, false); +} +EXPORT_SYMBOL(folio_clear_dirty_for_writethrough); + static void wb_inode_writeback_start(struct bdi_writeback *wb) { atomic_inc(&wb->writeback_inodes); --=20 2.55.0 From nobody Sat Oct 3 04:22:53 2026 Received: from mx0a-001b2d01.pphosted.com (mx0a-001b2d01.pphosted.com [148.163.156.1]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id A88F93DDB0A; Wed, 5 Aug 2026 06:29:45 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=148.163.156.1 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785911390; cv=none; b=Zg/cZHZ9z/dBuxzsiVMszOArqnOOMO6KMtKCOuIx+9QF/TI0oRryeGCuxManl1BPHlwMkQUZ2/MqpOYgPwXrgdSmWLOkcva7NJQbXuhW96cGYz7ZbhrlCeH/L6RReleFN0JqTg6/DlxKCX58FC1hWOtNGFG9iFUFuIkMSdcJNTg= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785911390; c=relaxed/simple; bh=dE70m9vMxqcxDFLDQ1hzZVu8qRknlnqRPTnCxnH9WFY=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=BJ3+Gc8A452YW11NpI1BS6dP9h5GIc2vvAfXMTwgNi8kVDKNdNS7JHfZiamELxMuRgLHn8QnuoIozkNpiApA25B86xHOSBox+3Hp9GUzSICZRzgeOlOKbeOb1PKTlyP8qZiidiS+tlIBCr+X5dZ6P7XfIAQ5h7hhgDT15V+UvkE= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.ibm.com; spf=pass smtp.mailfrom=linux.ibm.com; dkim=pass (2048-bit key) header.d=ibm.com header.i=@ibm.com header.b=bBvr2fzV; arc=none smtp.client-ip=148.163.156.1 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.ibm.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.ibm.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=ibm.com header.i=@ibm.com header.b="bBvr2fzV" Received: from pps.filterd (m0356517.ppops.net [127.0.0.1]) by mx0a-001b2d01.pphosted.com (8.18.1.11/8.18.1.11) with ESMTP id 6755lYV33005218; Wed, 5 Aug 2026 06:28:53 GMT DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=ibm.com; h=cc :content-transfer-encoding:date:from:in-reply-to:message-id :mime-version:references:subject:to; s=pp1; bh=VzXHAc99PQ8Pj4WEG R+XDuOXowi9q0bk7CdO8SzPF5w=; b=bBvr2fzVxG7VA/gr6N5kW73fFSkUZdy3w ES37ze7Sh6oezBRZfNTaddg4bfVNMGXNzD9R3TrR6yRRCIORrN0lHVdkAXo9z4Wl qGX53qlIfDlsEFUq5txbMLPJOEzdSn0acMmDHq3e014kplWkAv8HsRorA34vhZQ6 PXMMyZfjeQSdd3W5aclT80vfTj59k3xX+KBBvzOhiUCS2r5SrmwPtgoYxDK1Y37X km7V4TOJxed2zY5kotIowyDXTW5Cfh9ayesLb1r8wH9rsl6uqhkUWkk8Hbyf1Ypo a06Jkk9r/GBKWp5YU2wzSuYg2txsKx7gqjDVTe+M9Zd6FsgQaCuBA== Received: from ppma22.wdc07v.mail.ibm.com (5c.69.3da9.ip4.static.sl-reverse.com [169.61.105.92]) by mx0a-001b2d01.pphosted.com (PPS) with ESMTPS id 4fs8h51n5j-1 (version=TLSv1.2 cipher=ECDHE-RSA-AES256-GCM-SHA384 bits=256 verify=NOT); Wed, 05 Aug 2026 06:28:52 +0000 (GMT) Received: from pps.filterd (ppma22.wdc07v.mail.ibm.com [127.0.0.1]) by ppma22.wdc07v.mail.ibm.com (8.18.1.7/8.18.1.7) with ESMTP id 6756QGAE010624; Wed, 5 Aug 2026 06:28:51 GMT Received: from smtprelay01.fra02v.mail.ibm.com ([9.218.2.227]) by ppma22.wdc07v.mail.ibm.com (PPS) with ESMTPS id 4fsugw5efh-1 (version=TLSv1.2 cipher=ECDHE-RSA-AES256-GCM-SHA384 bits=256 verify=NOT); Wed, 05 Aug 2026 06:28:51 +0000 (GMT) Received: from smtpav07.fra02v.mail.ibm.com (smtpav07.fra02v.mail.ibm.com [10.20.54.106]) by smtprelay01.fra02v.mail.ibm.com (8.14.9/8.14.9/NCO v10.0) with ESMTP id 6756SnTn28180854 (version=TLSv1/SSLv3 cipher=DHE-RSA-AES256-GCM-SHA384 bits=256 verify=OK); Wed, 5 Aug 2026 06:28:49 GMT Received: from smtpav07.fra02v.mail.ibm.com (unknown [127.0.0.1]) by IMSVA (Postfix) with ESMTP id 049582004B; Wed, 5 Aug 2026 06:28:49 +0000 (GMT) Received: from smtpav07.fra02v.mail.ibm.com (unknown [127.0.0.1]) by IMSVA (Postfix) with ESMTP id C5CA620040; Wed, 5 Aug 2026 06:28:44 +0000 (GMT) Received: from li-dc0c254c-257c-11b2-a85c-98b6c1322444.ibm.com (unknown [9.124.211.239]) by smtpav07.fra02v.mail.ibm.com (Postfix) with ESMTP; Wed, 5 Aug 2026 06:28:44 +0000 (GMT) From: Ojaswin Mujoo To: Christian Brauner , linux-fsdevel@vger.kernel.org Cc: "Darrick J . Wong" , Carlos Maiolino , Alexander Viro , Jan Kara , Matthew Wilcox , Andrew Morton , Ritesh Harjani , Zhang Yi , Christoph Hellwig , Dave Chinner , Daniel Gomez , Pankaj Raghav , Theodore Tso , linux-xfs@vger.kernel.org, linux-kernel@vger.kernel.org, linux-mm@kvack.org Subject: [RFC PATCH v3 05/11] xfs: Add RWF_WRITETHROUGH support to xfs Date: Wed, 5 Aug 2026 11:58:11 +0530 Message-ID: X-Mailer: git-send-email 2.55.0 In-Reply-To: References: Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-TM-AS-GCONF: 00 X-Proofpoint-Reinject: loops=2 maxloops=12 X-Proofpoint-Spam-Info: AW1haW4tMjYwODA1MDA0NyBTYWx0ZWRfXyUXErExjDlHp OAzHyx0JKN6qnQycvqT/PyoGvF9V7PEaZz4h4Rcmaw+gYn9mtrAKNx16FIk9QAezB/qIt/RLUCR 0KYjQb2r7nt3i76wT8kWCVm6f+hAfe8= X-Proofpoint-Spam-Details-Enc: AW1haW4tMjYwODA1MDA0NyBTYWx0ZWRfXx2mb1u2yBckF QvBCvASeyp1TFd1xJ91Jg5x/aDeBvE6+lTTUxH9uoZ6A8SX00EsAQJC4/7AXNje1DZo2LsFV215 vcfU/w8rnJV9qKiRtesVsSjYBYtRGqLHiYdMRSwTKSvioPwgTK6XJiXce72pNU1D5MskyJm9/vV ixFWSxdUYt+2unXDRn9eLdAGZTd2Xkj+r/w5cgWnV+luRGB1wKvnSRbWbXIwOQW+Vv3l8bh1CBg 68VLO+s9oxvoBUbZABec3eAj2GmceULkTOoUEdUjWuU+z8558Bpo7cmSXc4K27wkVIhnNRst8M6 cb56tJlPrLlLnuO4By1f6YNIXWNzPrzlN3GfNwqYz6MVoch3U4fv6KTUEQmAMXGa0hpSmhvZhCC IiL0mz2hA8lORb5tneGJ+Dl8PFn72PSfzh0oyPuv8L160btCnOquDboV/maPXTmEgO2fqjBW5dl otHKReNq9jXZFzde/Cw== X-Authority-Analysis: v=2.4 cv=SI1ykuvH c=1 sm=1 tr=0 ts=6a72d825 cx=c_pps a=5BHTudwdYE3Te8bg5FgnPg==:117 a=5BHTudwdYE3Te8bg5FgnPg==:17 a=Sv0fKeRqtYgA:10 a=VkNPw1HP01LnGYTKEx00:22 a=RnoormkPH1_aCDwRdu11:22 a=U7nrCbtTmkRpXpFmAIza:22 a=pGLkceISAAAA:8 a=VnNF1IyMAAAA:8 a=FqVmevUCyaZ_8sojNSgA:9 X-Proofpoint-ORIG-GUID: zKijQG-cOKQ-oHjO_FM0F1bT9Hs5a1-B X-Proofpoint-GUID: 3uEF_eBIjDI1ffaCKDAAxBebhYLDgGk0 X-Proofpoint-Virus-Version: vendor=baseguard engine=ICAP:2.0.293,Aquarius:18.0.1176,Hydra:6.1.134,FMLib:17.12.100.49 definitions=2026-08-05_02,2026-08-04_02,2025-10-01_01 X-Proofpoint-Spam-Details: rule=outbound_notspam policy=outbound score=0 clxscore=1011 bulkscore=0 suspectscore=0 impostorscore=0 spamscore=0 phishscore=0 priorityscore=1501 lowpriorityscore=0 adultscore=0 malwarescore=0 classifier=typeunknown authscore=0 authtc= authcc= route=outbound adjust=0 reason=mlx scancount=1 engine=8.22.0-2606150000 definitions=main-2608050047 Content-Type: text/plain; charset="utf-8" Add the boilerplate needed to start supporting RWF_WRITETHROUGH in XFS. We use the direct write ->iomap_begin() functions to ensure the range under write through always has a real non-delalloc extent. We reuse the xfs dio's end IO function to perform extent conversion and i_size handling for us. *Note on COW extent over DATA hole case* In case of an unmapped COW extent over a DATA hole (due to COW preallocations), leave the extent unmapped until we are just about to send IO. At that time, use the ->writethrough_submit() call back to convert the COW extent to written. We initially tried converting during iomap begin() time (like dio does) but that results in a stale data exposure as follows: 1. iomap begin() - converts COW extent over DATA hole to written and marks IOMAP_F_NEW to handle zeroing. 2. During iomap_write_begin() -> realise extent is stale and return back without zeroing. 3. iomap begin() - Again sees the same COW extent but it's written this time so we don't mark IOMAP_F_NEW 4. Since IOMAP_F_NEW is unmarked, we never zeroout and hence expose stale data. To avoid the above, take the buffered IO approach of converting the extent just before IO, when we are sure to have zeroed out the folio. Co-developed-by: Ritesh Harjani (IBM) Signed-off-by: Ritesh Harjani (IBM) Signed-off-by: Ojaswin Mujoo --- fs/xfs/xfs_file.c | 83 +++++++++++++++++++++++++++++++++++++++++++---- 1 file changed, 77 insertions(+), 6 deletions(-) diff --git a/fs/xfs/xfs_file.c b/fs/xfs/xfs_file.c index 768cabf6250b..4b45ceacd461 100644 --- a/fs/xfs/xfs_file.c +++ b/fs/xfs/xfs_file.c @@ -702,6 +702,36 @@ static const struct iomap_dio_ops xfs_dio_write_ops = =3D { .end_io =3D xfs_dio_write_end_io, }; =20 +static int +xfs_writethrough_end_io( + struct iomap_writethrough_ctx *wt_ctx, + ssize_t size, + int error, + unsigned int flags) +{ + struct xfs_inode *ip =3D XFS_I(wt_ctx->inode); + xfs_off_t offset =3D wt_ctx->iocb->ki_pos; + + if (unlikely(error)) { + if (wt_ctx->flags & IOMAP_DIO_COW) + xfs_reflink_cancel_cow_range(ip, offset, size, true); + + return error; + } + + /* + * writethrough completions are handled same as dio with the exception + * that we need to explicitly change the i_disk_size. This is because + * unlike dio, we have already updated the i_size and hence the + * (i_disk_size < i_size) check will fail in dio code + */ + xfs_dio_write_end_io(wt_ctx->iocb, size, error, flags); + if (offset + size > ip->i_disk_size) + return xfs_setfilesize(ip, offset, size); + + return 0; +} + static void xfs_dio_zoned_submit_io( const struct iomap_iter *iter, @@ -1033,6 +1063,39 @@ xfs_file_dax_write( return ret; } =20 +static int +xfs_writethrough_submit( + struct inode *inode, + struct iomap *iomap, + loff_t offset, + u64 count) +{ + int error =3D 0; + unsigned int nofs_flag; + + /* + * Convert CoW extents to regular. + * + * We are under writethrough context with folio lock possibly held. To + * avoid memory allocation deadlocks, set the task-wide nofs context. + */ + if (iomap->flags & IOMAP_F_SHARED) { + nofs_flag =3D memalloc_nofs_save(); + error =3D xfs_reflink_convert_cow(XFS_I(inode), offset, count); + memalloc_nofs_restore(nofs_flag); + } + + return error; +} + +const struct iomap_writethrough_ops xfs_writethrough_ops =3D { + .ops =3D &xfs_direct_write_iomap_ops, + .write_ops =3D &xfs_iomap_write_ops, + .end_io =3D xfs_writethrough_end_io, + .writethrough_submit =3D &xfs_writethrough_submit +}; + + STATIC ssize_t xfs_file_buffered_write( struct kiocb *iocb, @@ -1055,9 +1118,13 @@ xfs_file_buffered_write( goto out; =20 trace_xfs_file_buffered_write(iocb, from); - ret =3D iomap_file_buffered_write(iocb, from, - &xfs_buffered_write_iomap_ops, &xfs_iomap_write_ops, - NULL); + if (iocb->ki_flags & IOCB_WRITETHROUGH) { + ret =3D iomap_file_writethrough_write(iocb, from, + &xfs_writethrough_ops, NULL); + } else + ret =3D iomap_file_buffered_write(iocb, from, + &xfs_buffered_write_iomap_ops, + &xfs_iomap_write_ops, NULL); =20 /* * If we hit a space limit, try to free up some lingering preallocated @@ -1092,8 +1159,12 @@ xfs_file_buffered_write( =20 if (ret > 0) { XFS_STATS_ADD(ip->i_mount, xs_write_bytes, ret); - /* Handle various SYNC-type writes */ - ret =3D generic_write_sync(iocb, ret); + /* + * Handle various SYNC-type writes. + * For writethrough, we handle sync during completion. + */ + if (!(iocb->ki_flags & IOCB_WRITETHROUGH)) + ret =3D generic_write_sync(iocb, ret); } return ret; } @@ -2104,7 +2175,7 @@ const struct file_operations xfs_file_operations =3D { .remap_file_range =3D xfs_file_remap_range, .fop_flags =3D FOP_MMAP_SYNC | FOP_BUFFER_RASYNC | FOP_BUFFER_WASYNC | FOP_DIO_PARALLEL_WRITE | - FOP_DONTCACHE, + FOP_DONTCACHE | FOP_WRITETHROUGH, .setlease =3D generic_setlease, }; =20 --=20 2.55.0 From nobody Sat Oct 3 04:22:53 2026 Received: from mx0a-001b2d01.pphosted.com (mx0a-001b2d01.pphosted.com [148.163.156.1]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 48E533DD87B; Wed, 5 Aug 2026 06:29:38 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=148.163.156.1 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785911383; cv=none; b=Tt8eA48AICew1TlzkX+bEWaTAHVOMVdblMnjoJkGetee9125njzNqB9g3eoQ7VRCqDw1gZm/pKloCjWmHK4XJ8yTsrHSSo6IEFpAs6kfJ614GVhBXtmjer6/Avj+GuqsoUVGErog/u/yP2ewYwMqZr8tCNb3L6uB9CrA249opoQ= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785911383; c=relaxed/simple; bh=v0R4+xsxsKsl1Oq5p/PdWg5SrI9vIKQWgK/dQyVj+k0=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=AvDdpuj2kt3bg9Th6lc+VK5DaQveWlvWwjAY8ZgwFLfnM0U1e2PxV133xLF5uZ+1mLAV8heTwWEqOWvM9/2km9m/bDgEHOh/w84cttpqvqdgs8idjN32JlPKA+lcxYn49KzXXPhvUN6F6lARpA0E96VlU4yVZSC7y8UhjgwPPfQ= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.ibm.com; spf=pass smtp.mailfrom=linux.ibm.com; dkim=pass (2048-bit key) header.d=ibm.com header.i=@ibm.com header.b=X5OUoKqw; arc=none smtp.client-ip=148.163.156.1 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.ibm.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.ibm.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=ibm.com header.i=@ibm.com header.b="X5OUoKqw" Received: from pps.filterd (m0356517.ppops.net [127.0.0.1]) by mx0a-001b2d01.pphosted.com (8.18.1.11/8.18.1.11) with ESMTP id 6755lYWs3005211; Wed, 5 Aug 2026 06:28:57 GMT DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=ibm.com; h=cc :content-transfer-encoding:date:from:in-reply-to:message-id :mime-version:references:subject:to; s=pp1; bh=89Qwln5dUFF/hnghQ LFiFdYuy9CZVGKkNR3U5GtxljQ=; b=X5OUoKqwIQdpiceBUc0iSJQCRBpIAf68I yGAEk51CgWfQ4f6L72iHrnoBrjDnfVPnDD4X3aziqYWjRX90M3bRb/DOTV6u8k+/ AZRmDPzpQsrZmszdQWwKN9jxF3I4i3PGDg2zn1pjOLY4i8sEum9v3VCC1AYHaK+W swnrujDoFHkqHVZnAr/f2ehQDkgEMEdxdBqEQB/rYP3FhMVQepGSQkk/uWVA5My+ IkYPZroq2Y2I9ghWPncualTyQgMHuQMEY2FUptn6OLsffYELUNXFo2+JK91Mva1r U1K8Y3QahB87y5G5l/4+zztKBmB01Uw3ChgtklURwIj2MqHHbs3JQ== Received: from ppma11.dal12v.mail.ibm.com (db.9e.1632.ip4.static.sl-reverse.com [50.22.158.219]) by mx0a-001b2d01.pphosted.com (PPS) with ESMTPS id 4fs8h51n66-1 (version=TLSv1.2 cipher=ECDHE-RSA-AES256-GCM-SHA384 bits=256 verify=NOT); Wed, 05 Aug 2026 06:28:57 +0000 (GMT) Received: from pps.filterd (ppma11.dal12v.mail.ibm.com [127.0.0.1]) by ppma11.dal12v.mail.ibm.com (8.18.1.7/8.18.1.7) with ESMTP id 6756QIXO001978; Wed, 5 Aug 2026 06:28:56 GMT Received: from smtprelay07.fra02v.mail.ibm.com ([9.218.2.229]) by ppma11.dal12v.mail.ibm.com (PPS) with ESMTPS id 4fswtyn3b4-1 (version=TLSv1.2 cipher=ECDHE-RSA-AES256-GCM-SHA384 bits=256 verify=NOT); Wed, 05 Aug 2026 06:28:56 +0000 (GMT) Received: from smtpav07.fra02v.mail.ibm.com (smtpav07.fra02v.mail.ibm.com [10.20.54.106]) by smtprelay07.fra02v.mail.ibm.com (8.14.9/8.14.9/NCO v10.0) with ESMTP id 6756Sshi44630382 (version=TLSv1/SSLv3 cipher=DHE-RSA-AES256-GCM-SHA384 bits=256 verify=OK); Wed, 5 Aug 2026 06:28:54 GMT Received: from smtpav07.fra02v.mail.ibm.com (unknown [127.0.0.1]) by IMSVA (Postfix) with ESMTP id 3ACF62004D; Wed, 5 Aug 2026 06:28:54 +0000 (GMT) Received: from smtpav07.fra02v.mail.ibm.com (unknown [127.0.0.1]) by IMSVA (Postfix) with ESMTP id 99B3F20040; Wed, 5 Aug 2026 06:28:49 +0000 (GMT) Received: from li-dc0c254c-257c-11b2-a85c-98b6c1322444.ibm.com (unknown [9.124.211.239]) by smtpav07.fra02v.mail.ibm.com (Postfix) with ESMTP; Wed, 5 Aug 2026 06:28:49 +0000 (GMT) From: Ojaswin Mujoo To: Christian Brauner , linux-fsdevel@vger.kernel.org Cc: "Darrick J . Wong" , Carlos Maiolino , Alexander Viro , Jan Kara , Matthew Wilcox , Andrew Morton , Ritesh Harjani , Zhang Yi , Christoph Hellwig , Dave Chinner , Daniel Gomez , Pankaj Raghav , Theodore Tso , linux-xfs@vger.kernel.org, linux-kernel@vger.kernel.org, linux-mm@kvack.org Subject: [RFC PATCH v3 06/11] iomap: Add aio support to RWF_WRITETHROUGH Date: Wed, 5 Aug 2026 11:58:12 +0530 Message-ID: <04cfaa7cf0e704d156f5accf6fb3f44921473612.1785908600.git.ojaswin@linux.ibm.com> X-Mailer: git-send-email 2.55.0 In-Reply-To: References: Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-TM-AS-GCONF: 00 X-Proofpoint-Reinject: loops=2 maxloops=12 X-Proofpoint-Spam-Info: AW1haW4tMjYwODA1MDA0NyBTYWx0ZWRfX7kKh22eMCGH/ 7VCokRaExW58RJNGfgxtfvdVxhbvvqAem17eR+uo+kERSQGgyJooqB717vwRHWan0JbXE0ntGK7 KIfbKXmLoGV1G5aS1x23HCJtJTbZ52Q= X-Proofpoint-Spam-Details-Enc: AW1haW4tMjYwODA1MDA0NyBTYWx0ZWRfX/IcSK8aqzhG6 W7rW1GbFVNJ4riUWTy1iBxNweCJIxd9CeTeLDlJsDcUP8KGWhwJHw3y36WexGGJsC/SH654LK3x QMFaXblZZalBOAzdD4SwfSN4u3e5DMSXbY/7Eaf7fDK1rqNEl5yUpnEU5CedoFAcsTY9Kf2DRRQ vKsgA8H0A/nFmUzCJcKyQ6H3ZVpB0+Ln54bQhncPkgjKFsx/fZiD5eek94LGlIQJSQvSZ6YXQHo YLME3ByXawmjCW1/QBhs3k0L8IjYW2fapFea/TKhVs2cNZBD1Ktd1GUm0fSv7/rmV4SvRLLNulh AGq9YNRKwCqgAp6CV3Ey68D2lvIjuEcK8Hv51mvdrIbtjYtJ4yHlAlwANkNFDXHJU/ckxx6IaCR 8XD8uDjc4A+MtVWuEFoonff7ttRxy1YGacELe+v1O6fwqxwvrp9LJ49CYG8n155FjQ8UsmYTb4V dA2lbJnSPv+OI9yNQ2w== X-Authority-Analysis: v=2.4 cv=SI1ykuvH c=1 sm=1 tr=0 ts=6a72d829 cx=c_pps a=aDMHemPKRhS1OARIsFnwRA==:117 a=aDMHemPKRhS1OARIsFnwRA==:17 a=Sv0fKeRqtYgA:10 a=VkNPw1HP01LnGYTKEx00:22 a=RnoormkPH1_aCDwRdu11:22 a=U7nrCbtTmkRpXpFmAIza:22 a=pGLkceISAAAA:8 a=VnNF1IyMAAAA:8 a=NGqb3fQY2RvKY4TqIeQA:9 X-Proofpoint-ORIG-GUID: Ds9PeaMLnIiSzO7nln_jjkDhrwwC0XeE X-Proofpoint-GUID: b90vshZtcTaJfVzD0Lz5KUt3jm1InLyl X-Proofpoint-Virus-Version: vendor=baseguard engine=ICAP:2.0.293,Aquarius:18.0.1176,Hydra:6.1.134,FMLib:17.12.100.49 definitions=2026-08-05_02,2026-08-04_02,2025-10-01_01 X-Proofpoint-Spam-Details: rule=outbound_notspam policy=outbound score=0 clxscore=1015 bulkscore=0 suspectscore=0 impostorscore=0 spamscore=0 phishscore=0 priorityscore=1501 lowpriorityscore=0 adultscore=0 malwarescore=0 classifier=typeunknown authscore=0 authtc= authcc= route=outbound adjust=0 reason=mlx scancount=1 engine=8.22.0-2606150000 definitions=main-2608050047 Content-Type: text/plain; charset="utf-8" With aio the only thing we need to be careful of is that writethrough can be in progress even after dropping inode and folio lock. Due to this, we need a way to synchronise with other paths where stable write is not enough, example: 1. Truncate to 0 in xfs sets i_size =3D 0 before waiting for writeback to complete. In case of writethrough, the end io completion can again push the i_size to a non-zero value. 2. Dio reads might race with aio writethrough ->end_io() and read 0s if unwritten conversion is yet to happen. Hence use the dio begin/end as it gives us the required guarantees. Co-developed-by: Ritesh Harjani (IBM) Signed-off-by: Ritesh Harjani (IBM) Signed-off-by: Ojaswin Mujoo --- fs/iomap/buffered-io.c | 54 ++++++++++++++++++++++++++++++++++++------ include/linux/iomap.h | 11 +++++++-- 2 files changed, 56 insertions(+), 9 deletions(-) diff --git a/fs/iomap/buffered-io.c b/fs/iomap/buffered-io.c index 3178e8c0fa13..22e4252dff4d 100644 --- a/fs/iomap/buffered-io.c +++ b/fs/iomap/buffered-io.c @@ -1202,6 +1202,9 @@ static ssize_t iomap_writethrough_complete(struct iom= ap_writethrough_ctx *wt_ctx =20 mapping_dec_inflight_stable_writes(inode->i_mapping); =20 + if (wt_ctx->is_aio) + inode_dio_end(inode); + if (!ret) { ret =3D wt_ctx->written; iocb->ki_pos +=3D ret; @@ -1211,12 +1214,27 @@ static ssize_t iomap_writethrough_complete(struct i= omap_writethrough_ctx *wt_ctx return ret; } =20 +static void iomap_writethrough_complete_work(struct work_struct *work) +{ + struct iomap_writethrough_ctx *wt_ctx =3D + container_of(work, struct iomap_writethrough_ctx, aio_work); + struct kiocb *iocb =3D wt_ctx->iocb; + + iocb->ki_complete(iocb, iomap_writethrough_complete(wt_ctx)); +} + static void iomap_writethrough_done(struct iomap_writethrough_ctx *wt_ctx) { - struct task_struct *waiter =3D wt_ctx->waiter; + if (!wt_ctx->is_aio) { + struct task_struct *waiter =3D wt_ctx->waiter; =20 - WRITE_ONCE(wt_ctx->waiter, NULL); - blk_wake_io_task(waiter); + WRITE_ONCE(wt_ctx->waiter, NULL); + blk_wake_io_task(waiter); + return; + } + + INIT_WORK(&wt_ctx->aio_work, iomap_writethrough_complete_work); + queue_work(wt_ctx->inode->i_sb->s_dio_done_wq, &wt_ctx->aio_work); } =20 static void iomap_writethrough_bio_end_io(struct bio *bio) @@ -1729,9 +1747,6 @@ ssize_t iomap_file_writethrough_write(struct kiocb *i= ocb, struct iov_iter *i, if (iocb_is_dsync(iocb)) /* D_SYNC support not implemented yet */ return -EOPNOTSUPP; - if (!is_sync_kiocb(iocb)) - /* aio support not implemented yet */ - return -EOPNOTSUPP; =20 /* * +1 to max bvecs to account for unaligned write spanning multiple @@ -1750,11 +1765,33 @@ ssize_t iomap_file_writethrough_write(struct kiocb = *iocb, struct iov_iter *i, wt_ctx->end_io =3D wt_ops->end_io; wt_ctx->old_i_size =3D i_size_read(inode); wt_ctx->max_bvecs =3D max_bvecs; + wt_ctx->is_aio =3D !is_sync_kiocb(iocb); atomic_set(&wt_ctx->ref, 1); - wt_ctx->waiter =3D current; + + if (!wt_ctx->is_aio) + wt_ctx->waiter =3D current; + else + /* + * With aio, writethrough can be in progress even after dropping + * inode and folio lock. Due to this, we need a way to + * synchronise with other paths where stable write is not enough + * (example truncate). Hence use the dio begin/end as it gives + * us the required guarantees. + */ + inode_dio_begin(inode); =20 mapping_inc_inflight_stable_writes(inode->i_mapping); =20 + if (wt_ctx->is_aio && !inode->i_sb->s_dio_done_wq) { + ret =3D sb_init_dio_done_wq(inode->i_sb); + if (ret < 0) { + mapping_dec_inflight_stable_writes(inode->i_mapping); + inode_dio_end(inode); + kfree(wt_ctx); + return ret; + } + } + blk_start_plug(&plug); =20 while ((ret =3D iomap_iter(&iter, wt_ops->ops)) > 0) { @@ -1769,6 +1806,9 @@ ssize_t iomap_file_writethrough_write(struct kiocb *i= ocb, struct iov_iter *i, cmpxchg(&wt_ctx->error, 0, ret); =20 if (!atomic_dec_and_test(&wt_ctx->ref)) { + if (wt_ctx->is_aio) + return -EIOCBQUEUED; + for (;;) { set_current_state(TASK_UNINTERRUPTIBLE); if (!READ_ONCE(wt_ctx->waiter)) diff --git a/include/linux/iomap.h b/include/linux/iomap.h index 427a2763221c..7203c4d92170 100644 --- a/include/linux/iomap.h +++ b/include/linux/iomap.h @@ -572,9 +572,16 @@ struct iomap_writethrough_ctx { atomic_t ref; unsigned int flags; int error; + bool is_aio; + + union { + /* used during submission and for non-aio completion */ + struct task_struct *waiter; + + /* used during aio completion */ + struct work_struct aio_work; + }; =20 - /* used during submission and for non-aio completion */ - struct task_struct *waiter; int (*end_io)(struct iomap_writethrough_ctx *wt_ctx, ssize_t size, int error, unsigned int flags); =20 --=20 2.55.0 From nobody Sat Oct 3 04:22:53 2026 Received: from mx0a-001b2d01.pphosted.com (mx0a-001b2d01.pphosted.com [148.163.156.1]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 002153DD84F; Wed, 5 Aug 2026 06:29:38 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=148.163.156.1 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785911384; cv=none; b=jsAwrmFboWKeuET1yMYoSkmeh5ELW3hvVkvF8L/RZ9C4p9Itt5UDEWl1mDAqKA3mj2498Qclxpd4TOk3MMDmic7a5HTUYb7zCRTxbtXY5V0SSHs475JEa1l4xVUj1+B8hhUwxGHnrgk41z8FfQ9ZIr4bqoB+oFDTYLSh2bUqHz0= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785911384; c=relaxed/simple; bh=YaYULzfGxLPLh9xyuOlwwtCNlPgHA7h3JE68ARTg1Co=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=g4A1McPILgTFIhor7dOYfBfsW1mHJVFqqx1sjc/sYmBh/5JbLPKUpbW7Ks1Om0hdQyViZOIpIKS/xOeiARwwEB9dxJ0ZiyHqZjrpPNvpQ8kLlzL6QwjTETEadxmc16/HQN6ruD/76mfqdM5GWhZl1VrLwdid8UNugd63XobNQyA= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.ibm.com; spf=pass smtp.mailfrom=linux.ibm.com; dkim=pass (2048-bit key) header.d=ibm.com header.i=@ibm.com header.b=trn3GdC7; arc=none smtp.client-ip=148.163.156.1 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.ibm.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.ibm.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=ibm.com header.i=@ibm.com header.b="trn3GdC7" Received: from pps.filterd (m0353729.ppops.net [127.0.0.1]) by mx0a-001b2d01.pphosted.com (8.18.1.11/8.18.1.11) with ESMTP id 6755li2v1121515; Wed, 5 Aug 2026 06:29:03 GMT DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=ibm.com; h=cc :content-transfer-encoding:date:from:in-reply-to:message-id :mime-version:references:subject:to; s=pp1; bh=NdKhp705dEcznPJBn iSr00N+kfQqAjZ7LiUlc9r+98w=; b=trn3GdC7qC0zMJlKNJL6+VcCZ58D84lUd kYaxVtn3sjkvFM3d7K90kCC7eYxjMIuUaxh7deNL8PDSCJKF/WiNQofAzJAjBYdz jH72FXGhdhe4WQfFHGso+xdipUQtkxcj2t6ac4+U491riKR/54KRLCq8PYpfMgk7 YGxz6Eaa8lJj9VP5uOvZvRDSOvbsyMyVK0iEYnxduvcOBrcytwFZIRDmr3o5zdGO Wg5zjbpu1PuD/e/RwH6nVPjFCQW4Gp7trhoY4VJrKhCj2p68UpOlzZEQWRuRjQya WuspRSDwczIEonXZ3TL138if0CvAwMff1g4dmQcHBYQuKez+06OIA== Received: from ppma12.dal12v.mail.ibm.com (dc.9e.1632.ip4.static.sl-reverse.com [50.22.158.220]) by mx0a-001b2d01.pphosted.com (PPS) with ESMTPS id 4fs8fqsjc2-1 (version=TLSv1.2 cipher=ECDHE-RSA-AES256-GCM-SHA384 bits=256 verify=NOT); Wed, 05 Aug 2026 06:29:02 +0000 (GMT) Received: from pps.filterd (ppma12.dal12v.mail.ibm.com [127.0.0.1]) by ppma12.dal12v.mail.ibm.com (8.18.1.7/8.18.1.7) with ESMTP id 6756QF6C028117; Wed, 5 Aug 2026 06:29:01 GMT Received: from smtprelay05.fra02v.mail.ibm.com ([9.218.2.225]) by ppma12.dal12v.mail.ibm.com (PPS) with ESMTPS id 4fsu4qnffh-1 (version=TLSv1.2 cipher=ECDHE-RSA-AES256-GCM-SHA384 bits=256 verify=NOT); Wed, 05 Aug 2026 06:29:01 +0000 (GMT) Received: from smtpav07.fra02v.mail.ibm.com (smtpav07.fra02v.mail.ibm.com [10.20.54.106]) by smtprelay05.fra02v.mail.ibm.com (8.14.9/8.14.9/NCO v10.0) with ESMTP id 6756SxEi49545516 (version=TLSv1/SSLv3 cipher=DHE-RSA-AES256-GCM-SHA384 bits=256 verify=OK); Wed, 5 Aug 2026 06:28:59 GMT Received: from smtpav07.fra02v.mail.ibm.com (unknown [127.0.0.1]) by IMSVA (Postfix) with ESMTP id 703552004D; Wed, 5 Aug 2026 06:28:59 +0000 (GMT) Received: from smtpav07.fra02v.mail.ibm.com (unknown [127.0.0.1]) by IMSVA (Postfix) with ESMTP id B6E2620040; Wed, 5 Aug 2026 06:28:54 +0000 (GMT) Received: from li-dc0c254c-257c-11b2-a85c-98b6c1322444.ibm.com (unknown [9.124.211.239]) by smtpav07.fra02v.mail.ibm.com (Postfix) with ESMTP; Wed, 5 Aug 2026 06:28:54 +0000 (GMT) From: Ojaswin Mujoo To: Christian Brauner , linux-fsdevel@vger.kernel.org Cc: "Darrick J . Wong" , Carlos Maiolino , Alexander Viro , Jan Kara , Matthew Wilcox , Andrew Morton , Ritesh Harjani , Zhang Yi , Christoph Hellwig , Dave Chinner , Daniel Gomez , Pankaj Raghav , Theodore Tso , linux-xfs@vger.kernel.org, linux-kernel@vger.kernel.org, linux-mm@kvack.org, Dave Chinner Subject: [RFC PATCH v3 07/11] iomap: Add DSYNC support to RWF_WRITETHROUGH Date: Wed, 5 Aug 2026 11:58:13 +0530 Message-ID: X-Mailer: git-send-email 2.55.0 In-Reply-To: References: Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-TM-AS-GCONF: 00 X-Proofpoint-Reinject: loops=2 maxloops=12 X-Proofpoint-GUID: Cv0jrwULaJcF3eYfANKQUI1WYQv-IJQm X-Proofpoint-ORIG-GUID: jiTnd4ONBctT5UEZ0brA3-HUtetCYSMT X-Proofpoint-Spam-Info: AW1haW4tMjYwODA1MDA0NyBTYWx0ZWRfXx9sw/ipQ/+3n vgCCYpEdXwZZlBX3Pwxquf91APDEKm1mPnWP4lkKo3sWrIrWrBrdSxG4/jIU8hiLHEntEGDIw+y bV5KTXQayb4dBXoPqcqcBB94RlbwGlU= X-Authority-Analysis: v=2.4 cv=K8cS2SWI c=1 sm=1 tr=0 ts=6a72d82e cx=c_pps a=bLidbwmWQ0KltjZqbj+ezA==:117 a=bLidbwmWQ0KltjZqbj+ezA==:17 a=Sv0fKeRqtYgA:10 a=VkNPw1HP01LnGYTKEx00:22 a=RnoormkPH1_aCDwRdu11:22 a=uAbxVGIbfxUO_5tXvNgY:22 a=VwQbUJbxAAAA:8 a=pGLkceISAAAA:8 a=VnNF1IyMAAAA:8 a=nOap7FSNCsi5F2RPnIcA:9 X-Proofpoint-Spam-Details-Enc: AW1haW4tMjYwODA1MDA0NyBTYWx0ZWRfX7MZYEjYorxrY gNi/Nl3hUJmyFEbf2QQ7PFnpXhMG9GBEoGUQ5jC31MEYJsW+Yswmi+1cUb8nzXNSxc7MHR27VOl 4s3+jCXA7zuMnlzdeUdeJdJhx4bqkWJb8wXWTeWnmZUwr5cM7sOGZrUITap76WTh6Q53lHoiZkk 4IbZ23SPgjMteo64RvdZl9Vate0PEreMOfWDZ85SabrsUWc+aqFgM/o86GPCJY0ZR9XorZdhFwS etSLqOQIi93+DOB1reQnVjVdDJTxjsH5PzSI03iCaSnnHTGAN5EMr3KlIyNxTfvHFU+BYB4UDlS AvhTm4f+e143eSCNSCf00hZEfJPXGHsXYMjQvrEqxVqDKLMrO5bdEtYU/sMUJx9Zi6mHFp9u2dp gqct9koOCpfJQuTXu6DUCGUmCovegNWhIVzynFFLeTgZEH7XR5T9HMRrCTNALq4hc/00PmAKvfs JtvoyCgTymDZAasyOXg== X-Proofpoint-Virus-Version: vendor=baseguard engine=ICAP:2.0.293,Aquarius:18.0.1176,Hydra:6.1.134,FMLib:17.12.100.49 definitions=2026-08-05_02,2026-08-04_02,2025-10-01_01 X-Proofpoint-Spam-Details: rule=outbound_notspam policy=outbound score=0 clxscore=1015 spamscore=0 impostorscore=0 bulkscore=0 priorityscore=1501 lowpriorityscore=0 malwarescore=0 phishscore=0 suspectscore=0 adultscore=0 classifier=typeunknown authscore=0 authtc= authcc= route=outbound adjust=0 reason=mlx scancount=1 engine=8.22.0-2606150000 definitions=main-2608050047 Content-Type: text/plain; charset="utf-8" Add DSYNC support to writethrough buffered writes. Unlike the usual buffered writes where we call generic_write_sync() inline during the syscall path, for writethrough we instead sync the data during IO completion path, just like dio. This allows aio writethrough to be truly async where the syscall can return after IO submission and the sync can then be done asynchronously during IO completion time. Further, just like dio, we utilize the FUA optimization, if available, to avoid syncing the data for DSYNC operations. Suggested-by: Dave Chinner Co-developed-by: Ritesh Harjani (IBM) Signed-off-by: Ritesh Harjani (IBM) Signed-off-by: Ojaswin Mujoo --- fs/iomap/buffered-io.c | 34 +++++++++++++++++++++++++++++++--- include/linux/iomap.h | 1 + 2 files changed, 32 insertions(+), 3 deletions(-) diff --git a/fs/iomap/buffered-io.c b/fs/iomap/buffered-io.c index 22e4252dff4d..d16694dd9995 100644 --- a/fs/iomap/buffered-io.c +++ b/fs/iomap/buffered-io.c @@ -1208,6 +1208,14 @@ static ssize_t iomap_writethrough_complete(struct io= map_writethrough_ctx *wt_ctx if (!ret) { ret =3D wt_ctx->written; iocb->ki_pos +=3D ret; + + /* + * If this is a DSYNC write and we couldn't optimize it, make + * sure we push it to stable storage now that we've written + * data. + */ + if (iocb_is_dsync(wt_ctx->iocb) && !wt_ctx->use_fua) + ret =3D generic_write_sync(iocb, ret); } =20 kfree(wt_ctx); @@ -1269,6 +1277,9 @@ iomap_writethrough_submit_bio(struct iomap_writethrou= gh_ctx *wt_ctx, for (i =3D 0; i < wt_ctx->nr_bvecs; i++) len +=3D wt_ctx->bvec[i].bv_len; =20 + if (wt_ctx->use_fua) + opf |=3D REQ_FUA; + bio =3D bio_alloc(iomap->bdev, wt_ctx->nr_bvecs, opf, GFP_NOFS); bio->bi_iter.bi_sector =3D iomap_sector(iomap, wt_ctx->bio_pos); bio->bi_end_io =3D iomap_writethrough_bio_end_io; @@ -1405,6 +1416,19 @@ static int iomap_writethrough_iter(struct iomap_writ= ethrough_ctx *wt_ctx, if (iter->iomap.type =3D=3D IOMAP_INLINE) return -EINVAL; =20 + /* + * If we realise that cache flush is necessary (eg FUA is not present + * or we need metadata updates) then we turn off the optimization. + */ + if (wt_ctx->use_fua) { + if (iter->iomap.type !=3D IOMAP_MAPPED || + (iter->iomap.flags & + (IOMAP_F_NEW | IOMAP_F_SHARED | IOMAP_F_DIRTY)) || + (bdev_write_cache(iter->iomap.bdev) && + !bdev_fua(iter->iomap.bdev))) + wt_ctx->use_fua =3D false; + } + do { struct folio *folio; size_t offset; /* Offset into folio */ @@ -1744,9 +1768,6 @@ ssize_t iomap_file_writethrough_write(struct kiocb *i= ocb, struct iov_iter *i, return -EINVAL; if (iocb->ki_flags & (IOCB_DONTCACHE)) return -EINVAL; - if (iocb_is_dsync(iocb)) - /* D_SYNC support not implemented yet */ - return -EOPNOTSUPP; =20 /* * +1 to max bvecs to account for unaligned write spanning multiple @@ -1768,6 +1789,13 @@ ssize_t iomap_file_writethrough_write(struct kiocb *= iocb, struct iov_iter *i, wt_ctx->is_aio =3D !is_sync_kiocb(iocb); atomic_set(&wt_ctx->ref, 1); =20 + /* + * Similar to dio, we optimistically set use_fua=3Dtrue to avoid explicit + * sync. In case we later realise cache flush is needed we set it back + * to false. + */ + wt_ctx->use_fua =3D iocb_is_dsync(iocb) && !(iocb->ki_flags & IOCB_SYNC); + if (!wt_ctx->is_aio) wt_ctx->waiter =3D current; else diff --git a/include/linux/iomap.h b/include/linux/iomap.h index 7203c4d92170..ba510d02c508 100644 --- a/include/linux/iomap.h +++ b/include/linux/iomap.h @@ -573,6 +573,7 @@ struct iomap_writethrough_ctx { unsigned int flags; int error; bool is_aio; + bool use_fua; =20 union { /* used during submission and for non-aio completion */ --=20 2.55.0 From nobody Sat Oct 3 04:22:53 2026 Received: from mx0a-001b2d01.pphosted.com (mx0a-001b2d01.pphosted.com [148.163.156.1]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id F20F13DD50F; Wed, 5 Aug 2026 06:29:38 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=148.163.156.1 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785911387; cv=none; b=cvJm7MSmEezv2bNS6FYwsH1nArTauTB1V2Ph/PfAKnY9d6v3zfK2URLHl+50CmuBvONYNJ+dW3RYGmYevgr/Vv9mBneF/m6Xlg4ghEZqZhZS9WvC4DRNwI7LKj24rDD4czpLFkcEX0zcmSitBXcQg73BtpWk8JklRyXa3JZL2aU= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785911387; c=relaxed/simple; bh=buyQrNlOM7YAU4cGVkAhmNiAzT2i59GO9B5AYwXgWq4=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=O+6jw6J6MyAiFRRhjZ05y1jQjbT+mxbjeiFUcTQloTT0vlgYwP6YGxew319rveGKe/jWQUvojPqL8FG39icUgH9cAWp9uCsggcGgYleGy/i4Lcw+9IPdm4UtsMGVE8htMSd8+IaKAHVraZOiAjcaa5hPojBKkuQq9JiD/DvDcrU= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.ibm.com; spf=pass smtp.mailfrom=linux.ibm.com; dkim=pass (2048-bit key) header.d=ibm.com header.i=@ibm.com header.b=qkxHJ+K9; arc=none smtp.client-ip=148.163.156.1 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.ibm.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.ibm.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=ibm.com header.i=@ibm.com header.b="qkxHJ+K9" Received: from pps.filterd (m0356517.ppops.net [127.0.0.1]) by mx0a-001b2d01.pphosted.com (8.18.1.11/8.18.1.11) with ESMTP id 6755lYV53005218; Wed, 5 Aug 2026 06:29:08 GMT DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=ibm.com; h=cc :content-transfer-encoding:date:from:in-reply-to:message-id :mime-version:references:subject:to; s=pp1; bh=DN/XoLF4HFzRvm+Qt OkdCzd7I9o/ktGE4HbJkuHIwqg=; b=qkxHJ+K9a/qMbPkt5BHIO9hv5+447+B9T iIgWF/4sIN6ye2aBxtnZP01wgBx1DG+jWoBWXBwie4NgaxxDTFBce3Ka3mja2VeK Jmdq1zd6elU0JK4rgHheh5OWZkENdZH75q4T7nbCkymVAoVnDn/y7sukh4/XA6ZH pQUVuu3Hz6Bku5RFTYJtdP8LTG+OzabhADWS5Yz2AeKjGZPJnFKD2bvtJQqJB8I7 E654CM1VIPLGwVJUntnnhNIVdM+Kzg6BfjVAKPtbHMCbU8qMVD9tAuGxUJjyhUBH s4fuDw8kg689Zv6coU/D7iPl610E+Ip+GxGHKwZZwX0WX2prFxv6A== Received: from ppma12.dal12v.mail.ibm.com (dc.9e.1632.ip4.static.sl-reverse.com [50.22.158.220]) by mx0a-001b2d01.pphosted.com (PPS) with ESMTPS id 4fs8h51n7c-1 (version=TLSv1.2 cipher=ECDHE-RSA-AES256-GCM-SHA384 bits=256 verify=NOT); Wed, 05 Aug 2026 06:29:07 +0000 (GMT) Received: from pps.filterd (ppma12.dal12v.mail.ibm.com [127.0.0.1]) by ppma12.dal12v.mail.ibm.com (8.18.1.7/8.18.1.7) with ESMTP id 6756QGih028129; Wed, 5 Aug 2026 06:29:06 GMT Received: from smtprelay03.fra02v.mail.ibm.com ([9.218.2.224]) by ppma12.dal12v.mail.ibm.com (PPS) with ESMTPS id 4fsu4qnfg4-1 (version=TLSv1.2 cipher=ECDHE-RSA-AES256-GCM-SHA384 bits=256 verify=NOT); Wed, 05 Aug 2026 06:29:06 +0000 (GMT) Received: from smtpav07.fra02v.mail.ibm.com (smtpav07.fra02v.mail.ibm.com [10.20.54.106]) by smtprelay03.fra02v.mail.ibm.com (8.14.9/8.14.9/NCO v10.0) with ESMTP id 6756T42F35914236 (version=TLSv1/SSLv3 cipher=DHE-RSA-AES256-GCM-SHA384 bits=256 verify=OK); Wed, 5 Aug 2026 06:29:04 GMT Received: from smtpav07.fra02v.mail.ibm.com (unknown [127.0.0.1]) by IMSVA (Postfix) with ESMTP id 7844220040; Wed, 5 Aug 2026 06:29:04 +0000 (GMT) Received: from smtpav07.fra02v.mail.ibm.com (unknown [127.0.0.1]) by IMSVA (Postfix) with ESMTP id EC7A52004D; Wed, 5 Aug 2026 06:28:59 +0000 (GMT) Received: from li-dc0c254c-257c-11b2-a85c-98b6c1322444.ibm.com (unknown [9.124.211.239]) by smtpav07.fra02v.mail.ibm.com (Postfix) with ESMTP; Wed, 5 Aug 2026 06:28:59 +0000 (GMT) From: Ojaswin Mujoo To: Christian Brauner , linux-fsdevel@vger.kernel.org Cc: "Darrick J . Wong" , Carlos Maiolino , Alexander Viro , Jan Kara , Matthew Wilcox , Andrew Morton , Ritesh Harjani , Zhang Yi , Christoph Hellwig , Dave Chinner , Daniel Gomez , Pankaj Raghav , Theodore Tso , linux-xfs@vger.kernel.org, linux-kernel@vger.kernel.org, linux-mm@kvack.org Subject: [RFC PATCH v3 08/11] fs: Introduce RWF_NOSERIAL flag to indicate parallel reads/writes Date: Wed, 5 Aug 2026 11:58:14 +0530 Message-ID: <71903b5bb03b193f2ad2e407f02e6f223802b436.1785908600.git.ojaswin@linux.ibm.com> X-Mailer: git-send-email 2.55.0 In-Reply-To: References: Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-TM-AS-GCONF: 00 X-Proofpoint-Reinject: loops=2 maxloops=12 X-Proofpoint-Spam-Info: AW1haW4tMjYwODA1MDA0NyBTYWx0ZWRfX2xVTIX8znp3p RHuNnf/iWdwNrXAe6GTCGhVQVofE+r95MA+lREbIIdpkfTxuALIxCaPvCIm0HjZtT/4bHD2/Tk0 WNLZVdDPqYsScuemub/CgMCCp5uCSFY= X-Proofpoint-Spam-Details-Enc: AW1haW4tMjYwODA1MDA0NyBTYWx0ZWRfX6eTjDZ7aAnEM qy9uZvReT1wDTnHCdm77Nl4oLZzynp7vn3lH4p9fzAIbr3H4UjxkryhG8Hk76GCytuSenZyB56y 1pFt5WanclxgkW01SyabBmMiBOJODJtvvTmEq3zNYas1ydKEiB9OsAFcqnrkmSkJUOQ6W9aRfLJ EfvMaZrkr+OQsYYtol/yixzF+ShmYlw2rquR7fhGXwCzqppZ4J7eVQ/k+RGuZfsukakbpcvXB3d ivbna4i+OEOTbNjizhgSB8I536zvfqEUiKtr3T64xs/+Zjmw3KaX/EHh31BcStk3PgGcil8IHHu DAxWG/OY0eeMhZyP0vO3tbrnAtnAz+V/bd2mxkO5vxTqGwxn0YNCenEGILeXW/QClxjaQhfND6I fmwtSMcPYl8wcjOPnJNMNxIwQUDvDXdTrP6XMbc9uk4Znsn0bVmb+e8iMVfiQIA6P0K+ZxkQNxI RHNZLqCQRvt+1vlDR7w== X-Authority-Analysis: v=2.4 cv=SI1ykuvH c=1 sm=1 tr=0 ts=6a72d833 cx=c_pps a=bLidbwmWQ0KltjZqbj+ezA==:117 a=bLidbwmWQ0KltjZqbj+ezA==:17 a=Sv0fKeRqtYgA:10 a=VkNPw1HP01LnGYTKEx00:22 a=RnoormkPH1_aCDwRdu11:22 a=U7nrCbtTmkRpXpFmAIza:22 a=pGLkceISAAAA:8 a=VnNF1IyMAAAA:8 a=S1zlp_KtrcgS9WxVI_oA:9 X-Proofpoint-ORIG-GUID: 4VnvKSaig3Wkkcc4-0zH4m0R7qbypOAK X-Proofpoint-GUID: Y4XoatcJmv_g6eEGocKw_TDUsXyI4E9p X-Proofpoint-Virus-Version: vendor=baseguard engine=ICAP:2.0.293,Aquarius:18.0.1176,Hydra:6.1.134,FMLib:17.12.100.49 definitions=2026-08-05_02,2026-08-04_02,2025-10-01_01 X-Proofpoint-Spam-Details: rule=outbound_notspam policy=outbound score=0 clxscore=1015 bulkscore=0 suspectscore=0 impostorscore=0 spamscore=0 phishscore=0 priorityscore=1501 lowpriorityscore=0 adultscore=0 malwarescore=0 classifier=typeunknown authscore=0 authtc= authcc= route=outbound adjust=0 reason=mlx scancount=1 engine=8.22.0-2606150000 definitions=main-2608050047 Content-Type: text/plain; charset="utf-8" Introduce RWF_NOSERIAL flag to indicate that the application is okay with its reads and writes going in parallel to other read/writes. This flag will allow writes and read to go in parallel (for eg, under a shared lock) increasing performance at the cost of losing the (loosely implemented) POSIX guarantees wrt to R/W serialization that various filesystems provide. In this patch we just introduce the flag and in upcoming patches we will use it to implement parallel writes to increase performance of RWF_WRITETHROUGH Co-developed-by: Ritesh Harjani (IBM) Signed-off-by: Ritesh Harjani (IBM) Signed-off-by: Ojaswin Mujoo --- include/linux/fs.h | 10 ++++++++++ include/uapi/linux/fs.h | 6 +++++- 2 files changed, 15 insertions(+), 1 deletion(-) diff --git a/include/linux/fs.h b/include/linux/fs.h index bff09a5f90e8..685ffe8da6ea 100644 --- a/include/linux/fs.h +++ b/include/linux/fs.h @@ -345,6 +345,7 @@ struct readahead_control; #define IOCB_DONTCACHE (__force int) RWF_DONTCACHE #define IOCB_NOSIGNAL (__force int) RWF_NOSIGNAL #define IOCB_WRITETHROUGH (__force int) RWF_WRITETHROUGH +#define IOCB_NOSERIAL (__force int) RWF_NOSERIAL =20 /* non-RWF related bits - start at 16 */ #define IOCB_EVENTFD (1 << 16) @@ -3475,6 +3476,15 @@ static inline int kiocb_set_rw_flags(struct kiocb *k= i, rwf_t flags, if (IS_DAX(ki->ki_filp->f_mapping->host)) return -EOPNOTSUPP; } + + /* + * For now we don't allow users to pass the NOSERIAL flag + * TODO: Once the semantics of this flag are finalized we can lift this + * restriction. + */ + if (flags & IOCB_NOSERIAL) + return -EOPNOTSUPP; + kiocb_flags |=3D (__force int) (flags & RWF_SUPPORTED); if (flags & RWF_SYNC) kiocb_flags |=3D IOCB_DSYNC; diff --git a/include/uapi/linux/fs.h b/include/uapi/linux/fs.h index 9c8d91b926a7..cde8c9492da6 100644 --- a/include/uapi/linux/fs.h +++ b/include/uapi/linux/fs.h @@ -454,10 +454,14 @@ typedef int __bitwise __kernel_rwf_t; /* buffered IO that is asynchronously written through to disk after write = */ #define RWF_WRITETHROUGH ((__force __kernel_rwf_t)0x00000200) =20 +/* buffered IO writes that are non sequential because they use a shared lo= ck */ +#define RWF_NOSERIAL ((__force __kernel_rwf_t)0x00000400) + /* mask of flags supported by the kernel */ #define RWF_SUPPORTED (RWF_HIPRI | RWF_DSYNC | RWF_SYNC | RWF_NOWAIT |\ RWF_APPEND | RWF_NOAPPEND | RWF_ATOMIC |\ - RWF_DONTCACHE | RWF_NOSIGNAL | RWF_WRITETHROUGH) + RWF_DONTCACHE | RWF_NOSIGNAL | RWF_WRITETHROUGH |\ + RWF_NOSERIAL) =20 #define PROCFS_IOCTL_MAGIC 'f' =20 --=20 2.55.0 From nobody Sat Oct 3 04:22:53 2026 Received: from mx0b-001b2d01.pphosted.com (mx0b-001b2d01.pphosted.com [148.163.158.5]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id E53083DCDA3; Wed, 5 Aug 2026 06:29:46 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=148.163.158.5 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785911390; cv=none; b=jiXgNZ5eS+MLYaQLZBxc599Cwe1doD7lh43lt3iMAq5Q7uiCndXzozIPneFqNCOJcRgYDkPtTc15wGM14dCfFeT+FNLEofI/dbJ/25I2bhiXF+Cei/Rg7uatZi2pxTg6l+jgYaaeB1Qwn9MYfELtJj8XkWtV5Mrc5a3dC4Vl2g0= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785911390; c=relaxed/simple; bh=XP0PfCtMVU9GtZcv9NEVEawIx1V1OxFMsgl1GQCeLA8=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=pascX+YUHykJc2vR8X4luOnBInBIvs9bU4z/3pED7yQDF8huA4SFmuf243fVsP4A1xFDRHEmjsirPqbT39eGXr5Q4q4tpvZkwqVt1xeesbw4Co6YJFsfEtwIe+0mzWJkMaJAYYzAgbzJwqXlgewZtQR3fOzYIbP+FaQckx3JPCk= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.ibm.com; spf=pass smtp.mailfrom=linux.ibm.com; dkim=pass (2048-bit key) header.d=ibm.com header.i=@ibm.com header.b=m4KkAuK+; arc=none smtp.client-ip=148.163.158.5 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.ibm.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.ibm.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=ibm.com header.i=@ibm.com header.b="m4KkAuK+" Received: from pps.filterd (m0353725.ppops.net [127.0.0.1]) by mx0a-001b2d01.pphosted.com (8.18.1.11/8.18.1.11) with ESMTP id 6755n0cZ2911541; Wed, 5 Aug 2026 06:29:13 GMT DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=ibm.com; h=cc :content-transfer-encoding:content-type:date:from:in-reply-to :message-id:mime-version:references:subject:to; s=pp1; bh=RWT8aC k7wXj77exlm/RNCevEhNoqqekH+wF/8jAK3ZM=; b=m4KkAuK+/TxYkgWr+JPiKC tSWSRQNv8tCgiD0Go5u/ZzCu/XVBtWg1gGpzsU+3SQMHUO79DA1yt2ChkVqbZwZc yeOyr1+Wu9BPO9tmbu3Ct6ewBmjDyIvFAPiqG1E8zSAqPN+fqesBT6BGNu0gtTde Ee34g2H62RGS6EPHzZg2L1nSVdE74dAzlk1uI+5lre2v32rC5Yovl6SbU9lBfbcV dtHfhQXb1PlSg6hSYpKNNwUfftAgar3IprM5q6vt8t86ShErLgb/XaIvzw9CHkfW EaeVR6uwUr+6HISeR66W2Eg7voRh0IKCeYEyU93Xu8uBNWEh2Mj9Fh9vSMS2MAlw == Received: from ppma13.dal12v.mail.ibm.com (dd.9e.1632.ip4.static.sl-reverse.com [50.22.158.221]) by mx0a-001b2d01.pphosted.com (PPS) with ESMTPS id 4fs77g99t1-1 (version=TLSv1.2 cipher=ECDHE-RSA-AES256-GCM-SHA384 bits=256 verify=NOT); Wed, 05 Aug 2026 06:29:12 +0000 (GMT) Received: from pps.filterd (ppma13.dal12v.mail.ibm.com [127.0.0.1]) by ppma13.dal12v.mail.ibm.com (8.18.1.7/8.18.1.7) with ESMTP id 6756QJe3026819; Wed, 5 Aug 2026 06:29:11 GMT Received: from smtprelay02.fra02v.mail.ibm.com ([9.218.2.226]) by ppma13.dal12v.mail.ibm.com (PPS) with ESMTPS id 4fswbgd4rj-1 (version=TLSv1.2 cipher=ECDHE-RSA-AES256-GCM-SHA384 bits=256 verify=NOT); Wed, 05 Aug 2026 06:29:11 +0000 (GMT) Received: from smtpav07.fra02v.mail.ibm.com (smtpav07.fra02v.mail.ibm.com [10.20.54.106]) by smtprelay02.fra02v.mail.ibm.com (8.14.9/8.14.9/NCO v10.0) with ESMTP id 6756T92950135454 (version=TLSv1/SSLv3 cipher=DHE-RSA-AES256-GCM-SHA384 bits=256 verify=OK); Wed, 5 Aug 2026 06:29:09 GMT Received: from smtpav07.fra02v.mail.ibm.com (unknown [127.0.0.1]) by IMSVA (Postfix) with ESMTP id C3C1120043; Wed, 5 Aug 2026 06:29:09 +0000 (GMT) Received: from smtpav07.fra02v.mail.ibm.com (unknown [127.0.0.1]) by IMSVA (Postfix) with ESMTP id 18E9D20040; Wed, 5 Aug 2026 06:29:05 +0000 (GMT) Received: from li-dc0c254c-257c-11b2-a85c-98b6c1322444.ibm.com (unknown [9.124.211.239]) by smtpav07.fra02v.mail.ibm.com (Postfix) with ESMTP; Wed, 5 Aug 2026 06:29:04 +0000 (GMT) From: Ojaswin Mujoo To: Christian Brauner , linux-fsdevel@vger.kernel.org Cc: "Darrick J . Wong" , Carlos Maiolino , Alexander Viro , Jan Kara , Matthew Wilcox , Andrew Morton , Ritesh Harjani , Zhang Yi , Christoph Hellwig , Dave Chinner , Daniel Gomez , Pankaj Raghav , Theodore Tso , linux-xfs@vger.kernel.org, linux-kernel@vger.kernel.org, linux-mm@kvack.org Subject: [RFC PATCH v3 09/11] xfs: Implement RWF_NOSERIAL to parallelize RWF_WRITETHROUGH writes Date: Wed, 5 Aug 2026 11:58:15 +0530 Message-ID: <945bb8c88a121580cb07d0311c6742a2584ea6b2.1785908600.git.ojaswin@linux.ibm.com> X-Mailer: git-send-email 2.55.0 In-Reply-To: References: Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable X-TM-AS-GCONF: 00 X-Proofpoint-Reinject: loops=2 maxloops=12 X-Proofpoint-Spam-Details-Enc: AW1haW4tMjYwODA1MDA0NyBTYWx0ZWRfX+dYRnCLDKhvr cVev75tHfJWa2pAIU9naHbh2Chgs+f09RgOA/nDPoLhK9mE/YkCkaihttWaHY+hIKozSYZynWg6 UzucviPoSIoPGyRdslfzvq3pVWNBO+Vx8T/z8ZYNKuGnOFnpXmCejHQ9HaizPwrs9UVc4eNqXc7 Iu5teIzUvzlmbU8lIKEm0kOljBLgd1PYJ2NJn8fLRVaqx+Iwnp2VrYrKWQhT8gEXNRsgNQwRDh6 GTXbYNXQg5zJIgthtK/k68BCOwas6UyqWi4IdUqWlSf/lLvLWE9qePkOze06QKuTqqA/o5IwTTk QMGy7CZO7PJwh9rTVADWaHqZOKzu3TglqvdKYlBo8pfY1QYzm+Y1XZaKfguxerK3CdFpAiEcwZt QWl4VgHKnE4FH/nJuarT99CN99VuU9zOj/MfbTfiU83ks0mpa11a9oCV3Mjr+Pv6ZAAPpHlxW6R +3QLoGhrZKNxW9Nni2w== X-Authority-Analysis: v=2.4 cv=WIFPmHsR c=1 sm=1 tr=0 ts=6a72d839 cx=c_pps a=AfN7/Ok6k8XGzOShvHwTGQ==:117 a=AfN7/Ok6k8XGzOShvHwTGQ==:17 a=IkcTkHD0fZMA:10 a=Sv0fKeRqtYgA:10 a=VkNPw1HP01LnGYTKEx00:22 a=RnoormkPH1_aCDwRdu11:22 a=V8glGbnc2Ofi9Qvn3v5h:22 a=pGLkceISAAAA:8 a=VnNF1IyMAAAA:8 a=ohjDDRDfXa_An5IlpZcA:9 a=3ZKOabzyN94A:10 a=QEXdDO2ut3YA:10 X-Proofpoint-GUID: 5xmYib4BB9HAko2uCn8GKxJzlikdQjnH X-Proofpoint-ORIG-GUID: fk7oH1CTDqvVf8GBmrVbweaiXebuZ106 X-Proofpoint-Spam-Info: AW1haW4tMjYwODA1MDA0NyBTYWx0ZWRfX5fz3zMi4oL0m x9sLx7z2MOXV8NhuQKb77HlAkYPMQVpvEOzEPOp4eTbewqKpD59TSV2aF7XjH4u5BzlF4iAYZI1 QltBsP9+mcpySN93DXmhMsQKCVIu4C8= X-Proofpoint-Virus-Version: vendor=baseguard engine=ICAP:2.0.293,Aquarius:18.0.1176,Hydra:6.1.134,FMLib:17.12.100.49 definitions=2026-08-05_02,2026-08-04_02,2025-10-01_01 X-Proofpoint-Spam-Details: rule=outbound_notspam policy=outbound score=0 spamscore=0 lowpriorityscore=0 priorityscore=1501 phishscore=0 malwarescore=0 suspectscore=0 clxscore=1015 impostorscore=0 bulkscore=0 adultscore=0 classifier=typeunknown authscore=0 authtc= authcc= route=outbound adjust=0 reason=mlx scancount=1 engine=8.22.0-2606150000 definitions=main-2608050047 In xfs, buffered writethrough writes take an exclusive inode lock similar to regular buffered write path. However, since writethrough also submits the write under the inode lock, the increased critical section really hurts performance when we have parallel writers writing to the same file. To mitigate this, implement RWF_NOSERIAL flag which allows us to perform the write under a shared inode lock, instead of an exclusive lock. This gives significant performance boost to single file, multiple writer workloads at the cost of losing write-write serialization and read-write serialization guarantees which XFS has historically provided. Let's look at each of the guarantees and what changes with the NOSERIAL flag: Write-write guarantee: Image 2 writers trying to write the same 3 folios. One is write As to all 3 folios (denoted by AAA) and the other is writing BBB. Then under exclusive lock the final state of the 3 folios could only be either AAA or BBB. However, with NOSERIAL writes, we could end up with mixed data in the 3 folios, like AAB, ABA, BAA etc. Note that this mixing will always happen at the inter folio level, data contained within the same folio will not be mixed as it is protected by the folio lock. Read-write guarantee: In XFS, reads also take a shared lock, however they don't take a folio lock when copying from the folio to the user buffer. Due to the shared lock of read and exclusive lock of write, we are able to guarantee that the read always reads either completely old data or completely new data. However, with the NOSERIAL flag, writethrough will take a shared lock for writes, ie a read can race with a write which is still in middle of copying data to the folio, hence the read can read a mix of old and new data. Despite the above changes in behavior, there might be applications who would be okay to lose the guarantees because of the nature of their workloads for example, if they never have multiple readers/writes doing IO to the same range in the file. Such applications would benefit significantly by using NOSERIAL writes. Below are some fio performance numbers of RWF_WRITETHROUGH with and without NOSERIAL writes. ** Multiple writes, single file (Pure overwrites, DSYNC) ** Fio Workload: libaio, buffered randwrite, bs=3D4k, size=3D2G (pre written), O_DSYNC iodepth=3D32 numjobs baseline BW RWF_NOSERIAL BW =CE=94% 1 350 392 +12.0% 2 526 798 +51.7% 4 597 1591 +166.5% 8 630 1839 +191.9% 16 570 1836 +222.1% ** Multiple writes, single file (Preallocated file, sync_file_range) ** Fio Workload: libaio, buffered randwrite, bs=3D16k, size=3D5G (fallocated) sync_file_range=3Dwait_before,write:16 numjobs baseline BW RWF_NOSERIAL BW =CE=94% 1 1099 1279 +16.4% 2 1482 2456 +65.7% 4 2065 2479 +20.1% 8 1778 2500 +40.6% 16 1787 2508 +40.3% * Multiple writes, single file (Truncated file, DSYNC) * Fio Workload: libaio, buffered randwrite, bs=3D4k, size=3D6.5G (truncated), O_DSYNC iodepth=3D32 numjobs baseline BW RWF_NOSERIAL BW =CE=94% 1 78 80 +2.6% 2 97 88 -9.3% 4 100 96 -4.0% 8 109 97 -11.0% 16 106 107 +0.9% Co-developed-by: Ritesh Harjani (IBM) Signed-off-by: Ritesh Harjani (IBM) Signed-off-by: Ojaswin Mujoo --- fs/xfs/xfs_file.c | 54 ++++++++++++++++++++++++++++++++++++++++------ include/linux/fs.h | 7 ++++++ 2 files changed, 55 insertions(+), 6 deletions(-) diff --git a/fs/xfs/xfs_file.c b/fs/xfs/xfs_file.c index 4b45ceacd461..b79076c15c5f 100644 --- a/fs/xfs/xfs_file.c +++ b/fs/xfs/xfs_file.c @@ -520,6 +520,42 @@ xfs_file_write_checks( return kiocb_modified(iocb); } =20 +STATIC ssize_t +xfs_file_writethrough_checks( + struct kiocb *iocb, + struct iov_iter *from, + unsigned int *iolock, + struct xfs_zone_alloc_ctx *ac) +{ + struct inode *inode =3D iocb->ki_filp->f_mapping->host; + size_t isize =3D i_size_read(inode); + size_t count =3D iov_iter_count(from); + ssize_t error; + + error =3D xfs_file_write_checks(iocb, from, iolock, ac); + if (error < 0) + return error; + + if (*iolock =3D=3D XFS_IOLOCK_EXCL) + return 0; + + /* + * Extending IO needs exclusive lock for i_size change + */ + if (iocb->ki_pos > isize || iocb->ki_pos + count >=3D isize) + goto upgrade_excl; + + return 0; + +upgrade_excl: + xfs_iunlock(XFS_I(inode), *iolock); + *iolock =3D XFS_IOLOCK_EXCL; + error =3D xfs_ilock_iocb(iocb, *iolock); + if (error) + *iolock =3D 0; + return error; +} + static ssize_t xfs_zoned_write_space_reserve( struct xfs_mount *mp, @@ -1108,23 +1144,29 @@ xfs_file_buffered_write( unsigned int iolock; =20 write_retry: - iolock =3D XFS_IOLOCK_EXCL; + if (iocb->ki_flags & IOCB_NOSERIAL) + iolock =3D XFS_IOLOCK_SHARED; + else + iolock =3D XFS_IOLOCK_EXCL; ret =3D xfs_ilock_iocb(iocb, iolock); if (ret) return ret; =20 - ret =3D xfs_file_write_checks(iocb, from, &iolock, NULL); - if (ret) - goto out; - trace_xfs_file_buffered_write(iocb, from); if (iocb->ki_flags & IOCB_WRITETHROUGH) { + ret =3D xfs_file_writethrough_checks(iocb, from, &iolock, NULL); + if (ret) + goto out; ret =3D iomap_file_writethrough_write(iocb, from, &xfs_writethrough_ops, NULL); - } else + } else { + ret =3D xfs_file_write_checks(iocb, from, &iolock, NULL); + if (ret) + goto out; ret =3D iomap_file_buffered_write(iocb, from, &xfs_buffered_write_iomap_ops, &xfs_iomap_write_ops, NULL); + } =20 /* * If we hit a space limit, try to free up some lingering preallocated diff --git a/include/linux/fs.h b/include/linux/fs.h index 685ffe8da6ea..ed5144512835 100644 --- a/include/linux/fs.h +++ b/include/linux/fs.h @@ -3495,6 +3495,13 @@ static inline int kiocb_set_rw_flags(struct kiocb *k= i, rwf_t flags, ki->ki_flags &=3D ~IOCB_APPEND; } =20 + /* + * Writethrough implies non-serial writes ie writes can go + * parallelly. + */ + if (flags & RWF_WRITETHROUGH) + kiocb_flags |=3D IOCB_NOSERIAL; + ki->ki_flags |=3D kiocb_flags; return 0; } --=20 2.55.0 From nobody Sat Oct 3 04:22:53 2026 Received: from mx0b-001b2d01.pphosted.com (mx0b-001b2d01.pphosted.com [148.163.158.5]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 09DCF3DDDCD; Wed, 5 Aug 2026 06:29:49 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=148.163.158.5 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785911399; cv=none; b=jIi08UwAJ5UiWhlOoWrAHN+3AdYyUnBG9m0RI8oC63fs04Dmx+NmybsWpwToYWWPiMtzIh5DlGdP7IfoxrXvT/r0q/cR4+FKZltfAP9gKEKSTcvGz+xVBKZ1DF3V0+tysalFxS+WZn9deR7TmE7SaJLJ2qnMY7M809zWpwQs/zM= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785911399; c=relaxed/simple; bh=i+bmx0L3bAhVnnyCIjjbNz5uCjMEmBoG8XakvHhL6qg=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=myTSBdW5WsboCkdMWen0VeYLe4u1uXeuAbLzxP6uv/1dlyqvGK+/qcrPFuSj21GhAj0qrOpMqKrrj8KyqSoC+rGKR8QR3aeH736pDg42VsOvgv4JJJ+ApukGqMDMOFnxhzCBsUNM/a0bh1CcoJZFJQpAHopCFGvo3lnlICNV4Kg= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.ibm.com; spf=pass smtp.mailfrom=linux.ibm.com; dkim=pass (2048-bit key) header.d=ibm.com header.i=@ibm.com header.b=eA9vtZBy; arc=none smtp.client-ip=148.163.158.5 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.ibm.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.ibm.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=ibm.com header.i=@ibm.com header.b="eA9vtZBy" Received: from pps.filterd (m0360072.ppops.net [127.0.0.1]) by mx0a-001b2d01.pphosted.com (8.18.1.11/8.18.1.11) with ESMTP id 6755mUSp3057676; Wed, 5 Aug 2026 06:29:18 GMT DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=ibm.com; h=cc :content-transfer-encoding:date:from:in-reply-to:message-id :mime-version:references:subject:to; s=pp1; bh=Zz5TTy6EEa7tg69Wf qzIDGIlVymyhC4BQmeTiGjvCPc=; b=eA9vtZByIb4Hj2Jip4yiVieTx4oQeAa20 6Nf/y9V4bl7IMlfG0el4F6qk9ggBe/RlfC916q4IKhZQAPTMG7MSJpFYN0kSvU77 9WrgF5RaywOISoIaLMS5BsI31AbnrwAuHGXW/e0kztxjHKOhRwSkrf7bMjfG6/mm oJW3vtjuMrfJUuObrMN6Z/V4hl8K4w0vdMUFYzUFZhKwf6gzSC3b6T766MhNvBZK B0cbXl2XscsxqYewjB8rTineddwvAUGvFYj8sXfyvzhIsIDstfU7UdsMPjB/vZrh +skk0A5QmTwafXxEXE1RunUfuLS65BcUwZqjs6GJtLvcUl9QW9IzA== Received: from ppma21.wdc07v.mail.ibm.com (5b.69.3da9.ip4.static.sl-reverse.com [169.61.105.91]) by mx0a-001b2d01.pphosted.com (PPS) with ESMTPS id 4fs8eus4f3-1 (version=TLSv1.2 cipher=ECDHE-RSA-AES256-GCM-SHA384 bits=256 verify=NOT); Wed, 05 Aug 2026 06:29:17 +0000 (GMT) Received: from pps.filterd (ppma21.wdc07v.mail.ibm.com [127.0.0.1]) by ppma21.wdc07v.mail.ibm.com (8.18.1.7/8.18.1.7) with ESMTP id 6756QIYw032086; Wed, 5 Aug 2026 06:29:16 GMT Received: from smtprelay06.fra02v.mail.ibm.com ([9.218.2.230]) by ppma21.wdc07v.mail.ibm.com (PPS) with ESMTPS id 4fsv4k5ayv-1 (version=TLSv1.2 cipher=ECDHE-RSA-AES256-GCM-SHA384 bits=256 verify=NOT); Wed, 05 Aug 2026 06:29:16 +0000 (GMT) Received: from smtpav07.fra02v.mail.ibm.com (smtpav07.fra02v.mail.ibm.com [10.20.54.106]) by smtprelay06.fra02v.mail.ibm.com (8.14.9/8.14.9/NCO v10.0) with ESMTP id 6756TE9e31916372 (version=TLSv1/SSLv3 cipher=DHE-RSA-AES256-GCM-SHA384 bits=256 verify=OK); Wed, 5 Aug 2026 06:29:14 GMT Received: from smtpav07.fra02v.mail.ibm.com (unknown [127.0.0.1]) by IMSVA (Postfix) with ESMTP id 868122004D; Wed, 5 Aug 2026 06:29:14 +0000 (GMT) Received: from smtpav07.fra02v.mail.ibm.com (unknown [127.0.0.1]) by IMSVA (Postfix) with ESMTP id 3289D20040; Wed, 5 Aug 2026 06:29:10 +0000 (GMT) Received: from li-dc0c254c-257c-11b2-a85c-98b6c1322444.ibm.com (unknown [9.124.211.239]) by smtpav07.fra02v.mail.ibm.com (Postfix) with ESMTP; Wed, 5 Aug 2026 06:29:09 +0000 (GMT) From: Ojaswin Mujoo To: Christian Brauner , linux-fsdevel@vger.kernel.org Cc: "Darrick J . Wong" , Carlos Maiolino , Alexander Viro , Jan Kara , Matthew Wilcox , Andrew Morton , Ritesh Harjani , Zhang Yi , Christoph Hellwig , Dave Chinner , Daniel Gomez , Pankaj Raghav , Theodore Tso , linux-xfs@vger.kernel.org, linux-kernel@vger.kernel.org, linux-mm@kvack.org Subject: [RFC PATCH v3 10/11] iomap: Avoid folio dirtying in case of RWF_WRITETHROUGH Date: Wed, 5 Aug 2026 11:58:16 +0530 Message-ID: X-Mailer: git-send-email 2.55.0 In-Reply-To: References: Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-TM-AS-GCONF: 00 X-Proofpoint-Reinject: loops=2 maxloops=12 X-Proofpoint-ORIG-GUID: HKXEHg4sQSXRf1lFCxUZ2J0r5qRupu2i X-Proofpoint-Spam-Info: AW1haW4tMjYwODA1MDA0NyBTYWx0ZWRfXzU0Yek79UgMU MQVwRm5dCa+ACL1eNmYBUIedf3OOdeYmdsdXiCEJZuM3OEPQmSrd3P+xg1WWcZOurQvUb0ovuyx 2MaRAKaVvY5Lyq1ewn4d114vIKXbSlw= X-Proofpoint-Spam-Details-Enc: AW1haW4tMjYwODA1MDA0NyBTYWx0ZWRfX5x3aRStL73li 3nWwUK1e9hIvnR1zb1+amhth1BbWn8v8+hZkzQQkvs5HO008RPPaXe7E8eBTb/F1/7DUD6chSEO R1fkVXuglLSfA/R2KOwquhMCws2IEo+n+DfG/U4bHXuEqfUPkSrKmpSUs9CkWHeoMgU4T0/w0Te Ouv6gTpDJtyMFvIRby86pakoKqpRBbYS6I87lywo8DRzzJtgJPl/8gbkEVrT5xnp2js6zzZ/XFv gnuDAWgROR4mb3eKLZwXCls+yhzyqSOp2bp1Qs/n05x06gEZXBqntrw8JbzKHAPCHd8lk7dzooc H46y5uIEwPB9ooYV2See268iutZ8HcYC5+3jPYqsRQG7h2X6FQtqMX0ldwSEyNNj/CFM1XL64Cz iAe/aCOtWaAf/i55mVloyEZ5kK7z+rL/UOzQnPHAaquNIFJuhpMPWF/CxuzgjcuxR9E+qJkEjDJ qOWFn4CKHx1NCOOi+rQ== X-Proofpoint-GUID: eLqPA66YBny84cHtUH1UZG4FWFEvxOqf X-Authority-Analysis: v=2.4 cv=KfzidwYD c=1 sm=1 tr=0 ts=6a72d83e cx=c_pps a=GFwsV6G8L6GxiO2Y/PsHdQ==:117 a=GFwsV6G8L6GxiO2Y/PsHdQ==:17 a=Sv0fKeRqtYgA:10 a=VkNPw1HP01LnGYTKEx00:22 a=RnoormkPH1_aCDwRdu11:22 a=RzCfie-kr_QcCd8fBx8p:22 a=pGLkceISAAAA:8 a=VnNF1IyMAAAA:8 a=0PHH8OhPgAe37TY-rOQA:9 X-Proofpoint-Virus-Version: vendor=baseguard engine=ICAP:2.0.293,Aquarius:18.0.1176,Hydra:6.1.134,FMLib:17.12.100.49 definitions=2026-08-05_02,2026-08-04_02,2025-10-01_01 X-Proofpoint-Spam-Details: rule=outbound_notspam policy=outbound score=0 lowpriorityscore=0 bulkscore=0 impostorscore=0 suspectscore=0 malwarescore=0 adultscore=0 clxscore=1015 priorityscore=1501 phishscore=0 spamscore=0 classifier=typeunknown authscore=0 authtc= authcc= route=outbound adjust=0 reason=mlx scancount=1 engine=8.22.0-2606150000 definitions=main-2608050047 Content-Type: text/plain; charset="utf-8" RWF_WRITETHROUGH dirties the folio only to send it for IO immediately, in the same context. Due to this, we can optimize away the folio dirtying and clearing step usually seen in buffered IO. Althrough we can do away with most of the accounting there are a couple of counters we need to take care of which we do during IO submission/completion. Co-developed-by: Ritesh Harjani (IBM) Signed-off-by: Ritesh Harjani (IBM) Signed-off-by: Ojaswin Mujoo --- fs/iomap/buffered-io.c | 64 +++++++++++++++++++++++++++++++++-------- include/linux/pagemap.h | 1 + mm/filemap.c | 21 ++++++++++++++ 3 files changed, 74 insertions(+), 12 deletions(-) diff --git a/fs/iomap/buffered-io.c b/fs/iomap/buffered-io.c index d16694dd9995..0844361fe0f1 100644 --- a/fs/iomap/buffered-io.c +++ b/fs/iomap/buffered-io.c @@ -11,6 +11,9 @@ #include #include #include +#include +#include "linux/pagemap.h" +#include "linux/page-flags.h" #include "internal.h" #include "trace.h" =20 @@ -1162,6 +1165,34 @@ static bool iomap_write_end_inline(const struct ioma= p_iter *iter, return true; } =20 +/* + * __iomap_writethrough_end() is almost same as __iomap_write_end() but wi= th the difference + * that we don't mark folio dirty since we are about to issue it for IO an= yways. + * Consequently, most of the accounting is skipped. + */ +static bool __iomap_writethrough_end(struct inode *inode, loff_t pos, size= _t len, + size_t copied, struct folio *folio) +{ + flush_dcache_folio(folio); + + /* + * The blocks that were entirely written will now be up-to-date, so we + * don't have to worry about a read_folio reading them and overwriting a + * partial write. However, if we've encountered a short write and only + * partially written into a block, it will not be marked up-to-date, so a + * read_folio might come in and destroy our partial write. + * + * Do the simplest thing and just treat any short write to a + * non-uptodate page as a zero-length write, and force the caller to + * redo the whole thing. + */ + if (unlikely(copied < len && !folio_test_uptodate(folio))) + return false; + iomap_set_range_uptodate(folio, offset_in_folio(folio, pos), len); + return true; +} + + /* * Returns true if all copied bytes have been written to the pagecache, * otherwise return false. @@ -1183,7 +1214,10 @@ static bool iomap_write_end(struct iomap_iter *iter,= size_t len, size_t copied, return bh_written =3D=3D copied; } =20 - return __iomap_write_end(iter->inode, pos, len, copied, folio); + if (iter->flags & IOMAP_WRITETHROUGH) + return __iomap_writethrough_end(iter->inode, pos, len, copied, folio); + else + return __iomap_write_end(iter->inode, pos, len, copied, folio); } =20 static ssize_t iomap_writethrough_complete(struct iomap_writethrough_ctx *= wt_ctx) @@ -1254,7 +1288,7 @@ static void iomap_writethrough_bio_end_io(struct bio = *bio) cmpxchg(&wt_ctx->error, 0, blk_status_to_errno(bio->bi_status)); bio_for_each_folio_all(fi, bio) - folio_end_writeback(fi.folio); + folio_end_writethrough(fi.folio, wt_ctx->error); =20 bio_put(bio); if (atomic_dec_and_test(&wt_ctx->ref)) @@ -1299,9 +1333,11 @@ iomap_writethrough_submit_bio(struct iomap_writethro= ugh_ctx *wt_ctx, =20 /* * In case of error we still need the I/O completion to run so we can - * release references and end writeback on the folios. + * release references, handle accounting and end writeback on the + * folios. */ if (error) { + task_io_account_cancelled_write(len); bio->bi_status =3D errno_to_blk_status(error); bio_endio(bio); return error; @@ -1351,6 +1387,7 @@ static void iomap_folio_prepare_writethrough(struct f= olio *folio, size_t off, { bool fully_written; u64 zero =3D 0; + u64 tmp_off =3D off; =20 if (folio_test_writeback(folio)) folio_wait_writeback(folio); @@ -1359,17 +1396,20 @@ static void iomap_folio_prepare_writethrough(struct= folio *folio, size_t off, folio_mark_dirty(folio); =20 /* - * We might either write through the complete folio or a partial folio - * writethrough might result in all blocks becoming non-dirty, so we need= to - * check and mark the folio clean if that is the case. + * For writethrough, we don't mark the write range dirty but we still + * need clear the dirty range if someone else has dirtied it before. + * Further, if the clearing results in folio becoming completely clean, + * then we need to take care of accounting. */ - fully_written =3D (off =3D=3D 0 && len =3D=3D folio_size(folio)); - iomap_clear_range_dirty(folio, off, len); - if (fully_written || - !iomap_find_dirty_range(folio, &zero, folio_size(folio))) - folio_clear_dirty_for_writethrough(folio); + if (iomap_find_dirty_range(folio, &tmp_off, tmp_off + len)) { + iomap_clear_range_dirty(folio, off, len); =20 - folio_start_writeback(folio); + if (!iomap_find_dirty_range(folio, &zero, folio_size(folio))) + folio_clear_dirty_for_writethrough(folio); + } + + task_io_account_write(folio_nr_pages(folio) * PAGE_SIZE); + folio_test_set_writeback(folio); } =20 /** diff --git a/include/linux/pagemap.h b/include/linux/pagemap.h index b20e38cc0fa0..774a2e57a9e0 100644 --- a/include/linux/pagemap.h +++ b/include/linux/pagemap.h @@ -1257,6 +1257,7 @@ void folio_wait_writeback(struct folio *folio); int folio_wait_writeback_killable(struct folio *folio); void end_page_writeback(struct page *page); void folio_end_writeback(struct folio *folio); +void folio_end_writethrough(struct folio *folio, bool error); void folio_end_writeback_no_dropbehind(struct folio *folio); void folio_end_dropbehind(struct folio *folio); void folio_wait_stable(struct folio *folio); diff --git a/mm/filemap.c b/mm/filemap.c index 58eb9d240643..a1a5f8837e03 100644 --- a/mm/filemap.c +++ b/mm/filemap.c @@ -1695,6 +1695,27 @@ void folio_end_writeback(struct folio *folio) } EXPORT_SYMBOL(folio_end_writeback); =20 +/** + * folio_end_writethrough - End writethrough against a folio + * @folio: The folio. + * error: Was there an error in IO. + * + * Context: May be called from process or interrupt context. + */ +void folio_end_writethrough(struct folio *folio, bool error) +{ + long nr =3D folio_nr_pages(folio); + + VM_BUG_ON_FOLIO(!folio_test_writeback(folio), folio); + + if (!error) + node_stat_mod_folio(folio, NR_WRITTEN, nr); + + if (folio_xor_flags_has_waiters(folio, 1 << PG_writeback)) + folio_wake_bit(folio, PG_writeback); +} +EXPORT_SYMBOL(folio_end_writethrough); + /** * __folio_lock - Get a lock on the folio, assuming we need to sleep to ge= t it. * @folio: The folio to lock --=20 2.55.0 From nobody Sat Oct 3 04:22:53 2026 Received: from mx0a-001b2d01.pphosted.com (mx0a-001b2d01.pphosted.com [148.163.156.1]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id ADAA13DC4B3; Wed, 5 Aug 2026 06:31:02 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=148.163.156.1 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785911466; cv=none; b=G21jvS9oMe5ijoOq1kB9Q2pqjL+9d7U7HBeM/NlP18onB+tx+h9MFIIOpKCXPXmgh4upaeBxbr7fjO/xIZJyW34du93CrIYdT2QrAsI+j+80aJFSZnzv9wPeWBJbSm/xUcUiWNuG/tSs+HZIf0CI6SqqI5+8BSZoDYGGgZs4BLA= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785911466; c=relaxed/simple; bh=SmPTURlB2CNqSYfYUmO+BqlxCqbDAMxql061z0A4d7M=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=k1kcrOnFrAaJ1cCfAypsRHwWBmJwyN8WyjX23JR+WuzI9gNf9lrwtmqoBxTG093sKDevbznWy9iyd5CiYSWGQ/fvAb++S7+ZeBIDtwm+j9dMZwPzRVCoMfcV0/AfQ/DWjAH9Jc1OeTa5e/xl2yipe1Hpo2tTqLLP0MjNQNq89Hc= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.ibm.com; spf=pass smtp.mailfrom=linux.ibm.com; dkim=pass (2048-bit key) header.d=ibm.com header.i=@ibm.com header.b=WwXWcHE8; arc=none smtp.client-ip=148.163.156.1 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.ibm.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.ibm.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=ibm.com header.i=@ibm.com header.b="WwXWcHE8" Received: from pps.filterd (m0353729.ppops.net [127.0.0.1]) by mx0a-001b2d01.pphosted.com (8.18.1.11/8.18.1.11) with ESMTP id 6755lrxa1121596; Wed, 5 Aug 2026 06:29:23 GMT DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=ibm.com; h=cc :content-transfer-encoding:date:from:in-reply-to:message-id :mime-version:references:subject:to; s=pp1; bh=095QS7upVPuSHTC8a 3xiMOI0ni0TS77b5nrnsLPg5w0=; b=WwXWcHE8zersWTROIS/lWp2pNcCZ62BlJ q0/uX4Fkz7Rp9ayC6yyUoxSg+MO/PffWXmiAO4Mqx36Elm1JSlS5pSmSHVCd7uPE JIKVLjFmmwNgjFtdvygQhPB7uhhzKvRFmBcHcX8g+KPvXSpkJxh6/xvZh3hV/+MY BsZEpftbLw0Yhf4ydx2R4CJJsKPx1HvAnEIYp6t2Vqe6YL/IdIICDk08PrxsNIxd nqBDKDmYXOnJtynD+P2ykWEof5PSmv412su4auKt/aeme6QG8SEaIz8gGwWYj+nJ +XxmATjpdfJQnPXNVdGWzZ8MP11Dcje0ht9ncZQzb1LTsHdvQDumQ== Received: from ppma11.dal12v.mail.ibm.com (db.9e.1632.ip4.static.sl-reverse.com [50.22.158.219]) by mx0a-001b2d01.pphosted.com (PPS) with ESMTPS id 4fs8fqsjdq-1 (version=TLSv1.2 cipher=ECDHE-RSA-AES256-GCM-SHA384 bits=256 verify=NOT); Wed, 05 Aug 2026 06:29:22 +0000 (GMT) Received: from pps.filterd (ppma11.dal12v.mail.ibm.com [127.0.0.1]) by ppma11.dal12v.mail.ibm.com (8.18.1.7/8.18.1.7) with ESMTP id 6756QFeP001943; Wed, 5 Aug 2026 06:29:21 GMT Received: from smtprelay04.fra02v.mail.ibm.com ([9.218.2.228]) by ppma11.dal12v.mail.ibm.com (PPS) with ESMTPS id 4fswtyn3du-1 (version=TLSv1.2 cipher=ECDHE-RSA-AES256-GCM-SHA384 bits=256 verify=NOT); Wed, 05 Aug 2026 06:29:21 +0000 (GMT) Received: from smtpav07.fra02v.mail.ibm.com (smtpav07.fra02v.mail.ibm.com [10.20.54.106]) by smtprelay04.fra02v.mail.ibm.com (8.14.9/8.14.9/NCO v10.0) with ESMTP id 6756TJMZ16711962 (version=TLSv1/SSLv3 cipher=DHE-RSA-AES256-GCM-SHA384 bits=256 verify=OK); Wed, 5 Aug 2026 06:29:19 GMT Received: from smtpav07.fra02v.mail.ibm.com (unknown [127.0.0.1]) by IMSVA (Postfix) with ESMTP id 682D620043; Wed, 5 Aug 2026 06:29:19 +0000 (GMT) Received: from smtpav07.fra02v.mail.ibm.com (unknown [127.0.0.1]) by IMSVA (Postfix) with ESMTP id 0697020040; Wed, 5 Aug 2026 06:29:15 +0000 (GMT) Received: from li-dc0c254c-257c-11b2-a85c-98b6c1322444.ibm.com (unknown [9.124.211.239]) by smtpav07.fra02v.mail.ibm.com (Postfix) with ESMTP; Wed, 5 Aug 2026 06:29:14 +0000 (GMT) From: Ojaswin Mujoo To: Christian Brauner , linux-fsdevel@vger.kernel.org Cc: "Darrick J . Wong" , Carlos Maiolino , Alexander Viro , Jan Kara , Matthew Wilcox , Andrew Morton , Ritesh Harjani , Zhang Yi , Christoph Hellwig , Dave Chinner , Daniel Gomez , Pankaj Raghav , Theodore Tso , linux-xfs@vger.kernel.org, linux-kernel@vger.kernel.org, linux-mm@kvack.org Subject: [RFC PATCH v3 11/11] iomap: Handle deadlock due to repeating folios in RWF_WRITETHROUGH Date: Wed, 5 Aug 2026 11:58:17 +0530 Message-ID: <7eb663a29c18e40d0000422792543ab3512c5a7e.1785908600.git.ojaswin@linux.ibm.com> X-Mailer: git-send-email 2.55.0 In-Reply-To: References: Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-TM-AS-GCONF: 00 X-Proofpoint-Reinject: loops=2 maxloops=12 X-Proofpoint-GUID: SPXSXqyhPvFwPef82RKiMuNzwMeR5-Ch X-Proofpoint-ORIG-GUID: aFxx8t1NCBaaen8xHPZ-fdaFxQOl4vLd X-Proofpoint-Spam-Info: AW1haW4tMjYwODA1MDA0NyBTYWx0ZWRfX6eAuWzQHDpRr R2qOTHDtgR5OiyzdWVJgrGDxpk16QRR5kKFq56ulmWg9iVHHJ+74Yh9D57EP4Rns4NzQHbFjs7f EIOarrUvspL6qnID0G4pwERs1h4BiKs= X-Authority-Analysis: v=2.4 cv=K8cS2SWI c=1 sm=1 tr=0 ts=6a72d842 cx=c_pps a=aDMHemPKRhS1OARIsFnwRA==:117 a=aDMHemPKRhS1OARIsFnwRA==:17 a=Sv0fKeRqtYgA:10 a=VkNPw1HP01LnGYTKEx00:22 a=RnoormkPH1_aCDwRdu11:22 a=uAbxVGIbfxUO_5tXvNgY:22 a=pGLkceISAAAA:8 a=VnNF1IyMAAAA:8 a=2R8uCszOOhyw9zs8MNsA:9 X-Proofpoint-Spam-Details-Enc: AW1haW4tMjYwODA1MDA0NyBTYWx0ZWRfX/YKbZ8sa0jQK dN9yhXJyqVxKbEaEm7mNWtHA7fYNkN8GfdHYq6nIfrnbTkbujt0hKd5xatg+4Sux66N63vG7fW4 NsL5lQD/xcC+Fa0PLsr/xGJty2Qo0xqBA6nqjriMW4/PHnNXVxIFPotef1QN0OoWn94LfNM9PWH tjdOECD4nmfqmKsITMMcf64tFaDsrwsU5EtOvFaZH+iWLykbOLepw+qJAB6diPN2lsRf5ybrp+j XljEuYrP+pe2xa6Irn2lnkAvTX+hvz7k8VrMFQa0vqw9ild136hkSN6Sdde+Mqt5DQv5cRrbIro KdstbJQ1/1GmK4ZgNqCI419wpVTL/bzzndjdzPgamJ5Egl4D0OjsmVpEc9KP9D7ZSAOQNULeUGJ VjP+X8Bh/NMyEDokK+JanIRQNZMHKyfoaDJYhYY6vL5UATRjHUudX9k/EdOY5pc/emmbnM4u07I gQ63yecp/+kCySsbZWA== X-Proofpoint-Virus-Version: vendor=baseguard engine=ICAP:2.0.293,Aquarius:18.0.1176,Hydra:6.1.134,FMLib:17.12.100.49 definitions=2026-08-05_02,2026-08-04_02,2025-10-01_01 X-Proofpoint-Spam-Details: rule=outbound_notspam policy=outbound score=0 clxscore=1011 spamscore=0 impostorscore=0 bulkscore=0 priorityscore=1501 lowpriorityscore=0 malwarescore=0 phishscore=0 suspectscore=0 adultscore=0 classifier=typeunknown authscore=0 authtc= authcc= route=outbound adjust=0 reason=mlx scancount=1 engine=8.22.0-2606150000 definitions=main-2608050047 Content-Type: text/plain; charset="utf-8" In iomap_writethrough_iter() we might encounter repeating folios across multiple iterations. Repeating folios can occur if, example, copy_folio_from_iter_atomic() does a short copy due to userspace pages not faulted in. This is an issue because a previous loop might have started writeback on them but not yet issued the IO. In the next iteration trying to get the same folio with FGP_STABLE will result in a deadlock. Since repeating folios will always be encountered back to back, we can just use a simple cur !=3D prev check to detect them. Use this to avoid waiting for writeback or starting writeback on folios we have already processed. Note that in ->endio() we might end up calling folio_end_writethrough() twice on the same folio which can cause issues with folio_xor_flags_has_waiters(). For simplicity, just change the folio_xor_flags_has_waiters() call to an idempotent variant. Reported-by: Pankaj Raghav Co-developed-by: Ritesh Harjani (IBM) Signed-off-by: Ritesh Harjani (IBM) Signed-off-by: Ojaswin Mujoo --- fs/iomap/buffered-io.c | 37 ++++++++++++++++++++++++++++++------- mm/filemap.c | 3 ++- 2 files changed, 32 insertions(+), 8 deletions(-) diff --git a/fs/iomap/buffered-io.c b/fs/iomap/buffered-io.c index 0844361fe0f1..70c1565e5416 100644 --- a/fs/iomap/buffered-io.c +++ b/fs/iomap/buffered-io.c @@ -804,6 +804,13 @@ struct folio *iomap_get_folio(struct iomap_iter *iter,= loff_t pos, size_t len) { fgf_t fgp =3D FGP_WRITEBEGIN; =20 + /* + * For writethrough, we open code the FGP_STABLE logic directly in + * iomap_writhrethrough_iter() so disable it here.. See + * iomap_writethrough_iter() for details. + */ + if (iter->flags & IOMAP_WRITETHROUGH) + fgp &=3D ~FGP_STABLE; if (iter->flags & IOMAP_NOWAIT) fgp |=3D FGP_NOWAIT; if (iter->flags & IOMAP_DONTCACHE) @@ -1383,15 +1390,12 @@ iomap_writethrough_try_submit(struct iomap_writethr= ough_ctx *wt_ctx, * need to clear the master dirty bit. */ static void iomap_folio_prepare_writethrough(struct folio *folio, size_t o= ff, - size_t len) + size_t len, bool already_prepared) { bool fully_written; u64 zero =3D 0; u64 tmp_off =3D off; =20 - if (folio_test_writeback(folio)) - folio_wait_writeback(folio); - if (folio_mkclean(folio)) folio_mark_dirty(folio); =20 @@ -1409,7 +1413,8 @@ static void iomap_folio_prepare_writethrough(struct f= olio *folio, size_t off, } =20 task_io_account_write(folio_nr_pages(folio) * PAGE_SIZE); - folio_test_set_writeback(folio); + if (!already_prepared) + folio_test_set_writeback(folio); } =20 /** @@ -1426,6 +1431,17 @@ static void iomap_folio_prepare_writethrough(struct = folio *folio, size_t off, * Folio handling note: We might be writing through a partial folio so we = need * to be careful to not clear the folio dirty bit unless there are no dirt= y blocks * in the folio after the writethrough. + * + * **A corner case to be careful about** + * + * For writethrough, we open code the stable write behavior to handle the = case + * where we encounter a folio that we already started writeback on but hav= e not + * yet submitted. In that case we must not wait for writeback again to avo= id + * deadlocking. Repeating folios can occur if, example, + * copy_folio_from_iter_atomic() does a short copy due to userspace pages = not + * faulted in. Also, repeating folios will always be encountered back to b= ack so + * we can just use a simple cur !=3D prev check to detect them. + */ static int iomap_writethrough_iter(struct iomap_writethrough_ctx *wt_ctx, struct iomap_iter *iter, struct iov_iter *i, @@ -1439,6 +1455,7 @@ static int iomap_writethrough_iter(struct iomap_write= through_ctx *wt_ctx, size_t chunk =3D mapping_max_folio_size(mapping); unsigned int bdp_flags =3D (iter->flags & IOMAP_NOWAIT) ? BDP_ASYNC : 0; unsigned int bs =3D i_blocksize(iter->inode); + struct folio *prev_folio =3D NULL; =20 /* copied over based on how DIO handles these flags */ if (iter->iomap.type =3D=3D IOMAP_UNWRITTEN) @@ -1530,6 +1547,10 @@ static int iomap_writethrough_iter(struct iomap_writ= ethrough_ctx *wt_ctx, if (mapping_writably_mapped(mapping)) flush_dcache_folio(folio); =20 + /* Open coding stable write behavior, see comment on top. */ + if (prev_folio !=3D folio) + folio_wait_writeback(folio); + copied =3D copy_folio_from_iter_atomic(folio, offset, bytes, i); written =3D iomap_write_end(iter, bytes, copied, folio) ? copied : 0; @@ -1544,8 +1565,10 @@ static int iomap_writethrough_iter(struct iomap_writ= ethrough_ctx *wt_ctx, off_aligned =3D round_down(offset, bs); len_aligned =3D round_up(offset + written, bs) - off_aligned; =20 - iomap_folio_prepare_writethrough(folio, off_aligned, - len_aligned); + iomap_folio_prepare_writethrough( + folio, off_aligned, len_aligned, prev_folio =3D=3D folio); + + prev_folio =3D folio; =20 if (!wt_ctx->nr_bvecs) { wt_ctx->bio_pos =3D round_down(pos, bs); diff --git a/mm/filemap.c b/mm/filemap.c index a1a5f8837e03..fc3ed0619838 100644 --- a/mm/filemap.c +++ b/mm/filemap.c @@ -1711,7 +1711,8 @@ void folio_end_writethrough(struct folio *folio, bool= error) if (!error) node_stat_mod_folio(folio, NR_WRITTEN, nr); =20 - if (folio_xor_flags_has_waiters(folio, 1 << PG_writeback)) + folio_test_clear_writeback(folio); + if (folio_test_waiters(folio)) folio_wake_bit(folio, PG_writeback); } EXPORT_SYMBOL(folio_end_writethrough); --=20 2.55.0