Issues / #1555

#1555 dpct migration: \cudaStream_t ext_stream_ = nullptr\ becomes a captured queue, so a multi-device stage runs on the constructing device's queue

open · @demetree · 1 コメント · GitHub で見る

NVIDIA / CUDAModels & quantsWindows

本文

### `cudaStream_t ext_stream_ = nullptr` migrated to a queue, and "no stream" became "this device's in-order queue"

Not Windows-specific: this is the dpct migration of a null-stream sentinel, and it affects any backend with more
than one device. Found by TeppeiLan1104 while testing the SYCL port (PR #1390) on two GPUs.

`sycl/include/strata/core/verify.hpp`:

```c++
-    cudaStream_t ext_stream_ = nullptr;
+    dpct::queue_ptr ext_stream_ = &dpct::get_in_order_queue();   // what dpct emits
```

The pointer form is evaluated **when the Verifier is constructed**, so the member initialiser captures the in-order
queue of whichever device happens to be current at that moment - not "no stream". Two places then compare against it:

```c++
-    if (ext_stream_ != nullptr) cs_ = ext_stream_;          // CUDA: null stream means "the default stream"
+    if (ext_stream_ != &dpct::get_in_order_queue()) cs_ = ext_stream_;   // sycl/src/core/verify.cpp init()
```

With the initialiser changed and the test unchanged, the condition is *almost never* true on the constructing device
(it compares a pointer to itself), so the assignment usually does not happen - but on a stage constructed on device 0
and `init`'ed on device 1, the in-order queue pointer differs and `cs_` takes **device 0's** queue:

```
stage [17, 48):  cs_ = 0000:0d:00.0     sh_cs_ = 0000:b0:00.0     copy_ = 0000:b0:00.0
```

Two failures follow, depending on the shared-expert path:

- with the 0.1.40 shared-expert fork: `Cannot submit to a queue with a dependency from a graph that is associated with a
  different context` at the first window;
- with `STRATA_SH_STREAM=0`: `UR_RESULT_ERROR_DEVICE_LOST`.

The same initialiser also decides teardown, where `if (cs_ && cs_ != ext_stream_)` is the "this stream is ours to
destroy" test.

### What I would change

Keep the sentinel a sentinel. `nullptr` means "no explicit stream", which is what the CUDA source says, and the
migration should either preserve that or rewrite every comparison site at once:

```c++
-    dpct::queue_ptr ext_stream_ = &dpct::get_in_order_queue();
+    dpct::queue_ptr ext_stream_ = nullptr;                       // nullptr = none until set_stream()
```

```c++
-    if (ext_stream_ != &dpct::get_in_order_queue()) cs_ = ext_stream_;
+    if (ext_stream_ != nullptr) cs_ = ext_stream_;
```

and the same in the destructor. I have carried this correction in the SYCL port's `sycl/tools/fixups.py` so a
re-migration keeps it, but the migration output itself should not emit a captured queue for a null sentinel - that
will bite the next dpct consumer the same way.

PR #1390 has the port's side with the measurements. Happy to test anything on two devices if a second one turns up.

関連リンク

インストール・モデル・リリースへの站内リンク。