]> git.ipfire.org Git - thirdparty/linux.git/commitdiff
block: stop the timeout timer when releasing a never added disk
authorChao Shi <coshi036@gmail.com>
Mon, 27 Jul 2026 20:12:57 +0000 (16:12 -0400)
committerJens Axboe <axboe@kernel.dk>
Wed, 29 Jul 2026 11:23:20 +0000 (05:23 -0600)
disk_release() undoes blk_mq_init_allocated_queue() for a disk whose
probe failed before add_disk(), but it only calls blk_mq_exit_queue().
Nothing there stops q->timeout, and that timer rolls forward: it stays
pending until it next expires, not until the last request completes.
So if the driver issued any I/O before adding the disk, the
request_queue is freed while still linked into a timer wheel bucket.

Commit 6f8191fdf41d ("block: simplify disk shutdown") dropped the
blk_cleanup_queue() call that used to stop it.  __del_gendisk() and
blk_mq_destroy_queue() still do; only the probe failure path lost it.

nvme gets there because nvme_update_ns_info() submits Report Zones or
FDP io-mgmt-recv on ns->queue before the disk is added, so a later
failure - a concurrent reset setting NVME_CTRL_FROZEN, or
device_add_disk() failing - lands in put_disk() with the timer armed:

  BUG: KASAN: slab-use-after-free in detach_if_pending+0x30c/0x340
  Write of size 8 at addr ffff888004d71310 by task kworker/u8:2/37
   __timer_delete_sync+0x156/0x240 kernel/time/timer.c:1621
   blk_sync_queue+0x22/0x40 block/blk-core.c:222
   nvme_sync_queues+0x100/0x150 drivers/nvme/host/core.c:5362
   nvme_reset_work+0x138/0x930 drivers/nvme/host/pci.c:3264

  Allocated by task 34:
   __blk_mq_alloc_disk+0x33/0x100 block/blk-mq.c:4462
   nvme_alloc_ns+0x290/0x3870 drivers/nvme/host/core.c:4146

  Freed by task 0:
   blk_free_queue_rcu+0x3a/0x50 block/blk-core.c:254
   rcu_core+0xc10/0x1730 kernel/rcu/tree.c:2857

The queue being synced there is ctrl->admin_q, only a victim sharing a
timer wheel bucket with the freed queue's dangling entry; other runs
tripped in enqueue_timer(), __run_timers() or blk_mq_timeout_work().
Failing nvme_alloc_ns() with a debug patch makes it deterministic: one
leaked timer trips KASAN within seconds, while 1987 patched releases
produced no splat.

Stop the timer and the queue work items before blk_mq_exit_queue(), like
blk_mq_destroy_queue() does.

Found by FuzzNvme.

Fixes: 6f8191fdf41d ("block: simplify disk shutdown")
Acked-by: Weidong Zhu <weizhu@fiu.edu>
Signed-off-by: Chao Shi <coshi036@gmail.com>
Reviewed-by: Christoph Hellwig <hch@lst.de>
Link: https://patch.msgid.link/20260727201257.211635-1-coshi036@gmail.com
Signed-off-by: Jens Axboe <axboe@kernel.dk>
block/genhd.c

index df2c3c69b467d47ff6aed269355a38fd7d7be33b..e8ce0cabf392ca531c9ff5443531904b0eb71720 100644 (file)
@@ -1281,14 +1281,18 @@ static void disk_release(struct device *dev)
        /*
         * To undo the all initialization from blk_mq_init_allocated_queue in
         * case of a probe failure where add_disk is never called we have to
-        * call blk_mq_exit_queue here. We can't do this for the more common
-        * teardown case (yet) as the tagset can be gone by the time the disk
-        * is released once it was added.
+        * call blk_mq_exit_queue here, after stopping the timer and work items
+        * that I/O issued before add_disk may have left pending.  We can't do
+        * this for the more common teardown case (yet) as the tagset can be
+        * gone by the time the disk is released once it was added.
         */
        if (queue_is_mq(disk->queue) &&
            test_bit(GD_OWNS_QUEUE, &disk->state) &&
-           !test_bit(GD_ADDED, &disk->state))
+           !test_bit(GD_ADDED, &disk->state)) {
+               blk_sync_queue(disk->queue);
+               blk_mq_cancel_work_sync(disk->queue);
                blk_mq_exit_queue(disk->queue);
+       }
 
        blkcg_exit_disk(disk);