[PATCH RFC 00/11] mm/migrate: separate migration paths and carry migration policy

Shivank Garg posted 11 patches 3 weeks, 2 days ago
Documentation/filesystems/locking.rst |   2 +-
Documentation/filesystems/vfs.rst     |   7 +-
fs/aio.c                              |   2 +-
fs/btrfs/disk-io.c                    |   5 +-
fs/btrfs/inode.c                      |   6 +-
fs/hugetlbfs/inode.c                  |   4 +-
fs/jfs/jfs_metapage.c                 |  14 +-
fs/nfs/internal.h                     |   2 +-
fs/nfs/write.c                        |   8 +-
include/linux/buffer_head.h           |   6 +-
include/linux/fs.h                    |   6 +-
include/linux/migrate.h               |   7 +-
include/linux/migrate_mode.h          |  14 +
include/linux/pagemap.h               |   2 +-
mm/compaction.c                       |   8 +-
mm/damon/ops-common.c                 |   7 +-
mm/gup.c                              |   7 +-
mm/memory-failure.c                   |   6 +-
mm/memory_hotplug.c                   |   6 +-
mm/mempolicy.c                        |  13 +-
mm/migrate.c                          | 717 ++++++++++++++++++++++------------
mm/page_alloc.c                       |   6 +-
mm/secretmem.c                        |   3 +-
mm/vmscan.c                           |   7 +-
virt/kvm/guest_memfd.c                |   2 +-
25 files changed, 563 insertions(+), 304 deletions(-)
[PATCH RFC 00/11] mm/migrate: separate migration paths and carry migration policy
Posted by Shivank Garg 3 weeks, 2 days ago
migrate_pages() handles hugetlb folios, movable_ops pages and LRU folios,
although their locking, mapping and retry requirements differ. It already
dispatches hugetlb folios to a dedicated engine but movable_ops 
pages still pass through the regular folio unmap and move machinery.

Migration policy has a separate interface problem. migrate_mode describes
blocking discipline, but it is passed independently from migrate_reason
through the core, and only migrate_mode reaches
address_space_operations->migrate_folio(). Adding another policy dimension
therefore requires more positional arguments, an overloaded mode or side
state. [3] 

This series separates the migration paths (classes) for movable_ops pages
and LRU folios, moves LRU policy into the LRU engine, and carries mode and
reason through struct migrate_control.

Thanks to Zi Yan and David Hildenbrand for the discussions and suggestions
that helped shape this series. I also used AI tools (Opus 5, GPT-5.6 Sol)
to refine my ideas, identify the edge-cases/bugs.

Old flow
========

migrate_pages(mode, reason)
  |
  +-- migrate_hugetlbs()
  |     `-- scan the mixed source list (from) on each retry
  |
  `-- split the remaining list into batches
        |
        +-- MIGRATE_ASYNC
        |     `-- migrate_pages_batch()
        |           inline unmap/retry -> flush -> move/retry
        |
        `-- MIGRATE_SYNC[_LIGHT]
              `-- migrate_pages_sync()
                   +-- try async batch pre-pass
                   `-- retry failures one at a time through
                       migrate_pages_batch()

movable_ops pages are identified by branches inside the batch unmap and
move paths.

New flow
========

migrate_pages(ctl)
  |
  +-- migrate_hugetlbs()
  |     `-- collect and retry hugetlb folios
  |
  +-- migrate_movable_ops_pages()
  |     `-- allocate, lock, invoke the movable_ops callback and retry
  |
  `-- migrate_lru_folios()
        `-- form bounded LRU batches
              `-- __migrate_lru_folios()
                    +-- async: migrate_folios_batch()
                    `-- sync:
                        +-- async batch pre-pass
                        `-- retry failures one at a time

migrate_folios_batch()
  |
  +-- migrate_folios_unmap()
  +-- try_to_unmap_flush()
  +-- migrate_folios_move()
  `-- migrate_folios_undo()

migrate_pages() now only dispatches classes. Each entry owns its lists and
retries; the LRU entry also owns batching and synchronous fallback.

Motivation
==========

This refactor makes the migration control flow easier to follow and gives
new optimizations clear boundaries and insertion points.

Separating movable_ops also prepares for memdescs, where these pages are
expected to lose their folio representation and folio->lru handoff. The
current list interface remains, but its folio dependencies are now isolated.

struct migrate_control is caller-owned, stack allocated and passed as const.
It carries mode and reason through migrate_pages() and ->migrate_folio(),
and provide place for future extension to policy. This is one-time pain of
updating callers, but doing it once avoids repeating that churn whenever
a new use case or optimization need additional policy.

Follow-on 
=========

This series establishes the class, phase and policy boundaries without
adding copy implementation. Possible follow-on use looks like:

Yiannis and Alirad's RFC adds an asynchronous non-temporal mode and updates
existing asynchronous-mode checks. [1] A separate policy field would keep
cache behavior independent of blocking discipline.

  /* migrate_pages() caller: demote_folio_list() */
  ctl.reason = MR_DEMOTION;   /* migration reason - already exists */
  ctl.mode = MIGRATE_ASYNC;   /* blocking discipline - already exists */
  ctl.cache_hint = MIGRATE_COPY_NT; /* caching intent- new policy */

  __migrate_folio(..., ctl);       /* NT/offload based on caller's hint*/
     `-> folio_mc_copy() or folio_mc_copy_nt() or ...

My batch-copy/offload series [2] can extend migrate_control to take caller's
preference of copy engines (like DMA offload, multi-threaded copy, etc.) or
batch size. Some of this may remain wishful thinking but that is what the
RFC is for :)

Behavior Changes
================

No functional change is intended for the LRU and hugetlb migration paths.

The movable_ops pass has three intentional differences:

  - movable_ops pages are attempted before LRU folios;
  - movable_ops pages no longer use the LRU asynchronous pre-pass.
    synchronous callers use their requested mode from the first attempt.
  - movable_ops callback returning -EAGAIN releases the destination, so
    retry allocates a new one.

Unmigrated pages are still returned through the original source list.

[1] https://lore.kernel.org/r/20260730-rfc-nt-demote-v2-0-452dbe3b5073@zptcorp.com
[2] https://lore.kernel.org/r/20260630-shivank-batch-migrate-offload-v6-0-da95d7e8b8a2@amd.com 
[3] https://lore.kernel.org/r/cae6ab98-3441-38bd-1e07-4586a85cdc74@google.com

---
Shivank Garg (11):
      mm/migrate: extract folio unmap phase
      mm/migrate: handle retries in migrate_folios_move()
      mm/migrate: factor out folio splitting on allocation failure
      mm/migrate: use a dedicated list for hugetlb folios
      mm/migrate: add a dedicated movable_ops migration pass
      mm/migrate: rename migrate_pages_batch() to migrate_folios_batch()
      mm/migrate: add migrate_lru_folios() entry point
      mm/migrate: move LRU batching into migrate_lru_folios()
      mm/migrate: thread migration policy through a control struct
      mm/migrate: pass migrate_control to migrate_pages()
      mm/migrate: pass migrate_control to migrate_folio()

 Documentation/filesystems/locking.rst |   2 +-
 Documentation/filesystems/vfs.rst     |   7 +-
 fs/aio.c                              |   2 +-
 fs/btrfs/disk-io.c                    |   5 +-
 fs/btrfs/inode.c                      |   6 +-
 fs/hugetlbfs/inode.c                  |   4 +-
 fs/jfs/jfs_metapage.c                 |  14 +-
 fs/nfs/internal.h                     |   2 +-
 fs/nfs/write.c                        |   8 +-
 include/linux/buffer_head.h           |   6 +-
 include/linux/fs.h                    |   6 +-
 include/linux/migrate.h               |   7 +-
 include/linux/migrate_mode.h          |  14 +
 include/linux/pagemap.h               |   2 +-
 mm/compaction.c                       |   8 +-
 mm/damon/ops-common.c                 |   7 +-
 mm/gup.c                              |   7 +-
 mm/memory-failure.c                   |   6 +-
 mm/memory_hotplug.c                   |   6 +-
 mm/mempolicy.c                        |  13 +-
 mm/migrate.c                          | 717 ++++++++++++++++++++++------------
 mm/page_alloc.c                       |   6 +-
 mm/secretmem.c                        |   3 +-
 mm/vmscan.c                           |   7 +-
 virt/kvm/guest_memfd.c                |   2 +-
 25 files changed, 563 insertions(+), 304 deletions(-)
---
base-commit: cee9395acd8043be0644b25c34bfa86623f2b935
change-id: 20260824-migrate-refactor-shivank-4ee3949fcdcb

Best regards,
-- 
Shivank Garg <shivankg@amd.com>