Re-run only the failing playbook on the same matrix entry by triggering the workflow with the playbook id:
ghts: 99%|█████████▉| 2114/2130 [19:25<00:04, 3.30it/s]
Loading weights: 100%|█████████▉| 2120/2130 [19:26<00:02, 3.83it/s]
Loading weights: 100%|█████████▉| 2125/2130 [19:28<00:01, 3.35it/s]
Loading weights: 100%|█████████▉| 2128/2130 [19:29<00:00, 3.28it/s]
Loading weights: 100%|██████████| 2130/2130 [19:30<00:00, 2.98it/s]
Loading weights: 100%|██████████| 2130/2130 [19:30<00:00, 1.82it/s]
Unsloth: Tokenizing ["text"] (num_proc=27): 0%| | 0/128 [00:00<?, ? examples/s]
Unsloth: Tokenizing ["text"] (num_proc=27): 4%|▍ | 5/128 [00:02<01:09, 1.76 examples/s]
Unsloth: Tokenizing ["text"] (num_proc=27): 8%|▊ | 10/128 [00:04<00:53, 2.22 examples/s]
Unsloth: Tokenizing ["text"] (num_proc=27): 12%|█▏ | 15/128 [00:06<00:46, 2.42 examples/s]
Unsloth: Tokenizing ["text"] (num_proc=27): 16%|█▌ | 20/128 [00:08<00:42, 2.53 examples/s]
Unsloth: Tokenizing ["text"] (num_proc=27): 20%|█▉ | 25/128 [00:10<00:39, 2.59 examples/s]
Unsloth: Tokenizing ["text"] (num_proc=27): 23%|██▎ | 30/128 [00:12<00:37, 2.62 examples/s]
Unsloth: Tokenizing ["text"] (num_proc=27): 27%|██▋ | 35/128 [00:13<00:34, 2.66 examples/s]
Unsloth: Tokenizing ["text"] (num_proc=27): 31%|███▏ | 40/128 [00:15<00:32, 2.68 examples/s]
Unsloth: Tokenizing ["text"] (num_proc=27): 35%|███▌ | 45/128 [00:17<00:30, 2.69 examples/s]
Unsloth: Tokenizing ["text"] (num_proc=27): 39%|███▉ | 50/128 [00:19<00:28, 2.70 examples/s]
Unsloth: Tokenizing ["text"] (num_proc=27): 43%|████▎ | 55/128 [00:21<00:27, 2.69 examples/s]
Unsloth: Tokenizing ["text"] (num_proc=27): 47%|████▋ | 60/128 [00:23<00:25, 2.70 examples/s]
Unsloth: Tokenizing ["text"] (num_proc=27): 51%|█████ | 65/128 [00:24<00:23, 2.71 examples/s]
Unsloth: Tokenizing ["text"] (num_proc=27): 55%|█████▍ | 70/128 [00:26<00:21, 2.71 examples/s]
Unsloth: Tokenizing ["text"] (num_proc=27): 59%|█████▊ | 75/128 [00:28<00:19, 2.71 examples/s]
Unsloth: Tokenizing ["text"] (num_proc=27): 62%|██████▎ | 80/128 [00:30<00:17, 2.72 examples/s]
Unsloth: Tokenizing ["text"] (num_proc=27): 66%|██████▋ | 85/128 [00:32<00:15, 2.71 examples/s]
Unsloth: Tokenizing ["text"] (num_proc=27): 70%|███████ | 90/128 [00:34<00:13, 2.72 examples/s]
Unsloth: Tokenizing ["text"] (num_proc=27): 74%|███████▍ | 95/128 [00:35<00:12, 2.74 examples/s]
Unsloth: Tokenizing ["text"] (num_proc=27): 78%|███████▊ | 100/128 [00:37<00:10, 2.74 examples/s]
Unsloth: Tokenizing ["text"] (num_proc=27): 81%|████████▏ | 104/128 [00:39<00:09, 2.57 examples/s]
Unsloth: Tokenizing ["text"] (num_proc=27): 84%|████████▍ | 108/128 [00:41<00:08, 2.46 examples/s]
Unsloth: Tokenizing ["text"] (num_proc=27): 88%|████████▊ | 112/128 [00:43<00:06, 2.37 examples/s]
Unsloth: Tokenizing ["text"] (num_proc=27): 91%|█████████ | 116/128 [00:45<00:05, 2.31 examples/s]
Unsloth: Tokenizing ["text"] (num_proc=27): 94%|█████████▍| 120/128 [00:46<00:03, 2.26 examples/s]
Unsloth: Tokenizing ["text"] (num_proc=27): 97%|█████████▋| 124/128 [00:48<00:01, 2.23 examples/s]
Unsloth: Tokenizing ["text"] (num_proc=27): 100%|██████████| 128/128 [00:50<00:00, 2.22 examples/s]
Unsloth: Tokenizing ["text"] (num_proc=27): 100%|██████████| 128/128 [00:50<00:00, 2.51 examples/s]
Map: 0%| | 0/128 [00:00<?, ? examples/s]
Map: 100%|██████████| 128/128 [00:00<00:00, 5718.94 examples/s]
The tokenizer has new PAD/BOS/EOS tokens that differ from the model config and generation config. The model config and generation config were aligned accordingly, being updated with the tokenizer's values. Updated tokens: {'bos_token_id': 2}.
[transformers.trainer_utils|WARNING]The tokenizer has new PAD/BOS/EOS tokens that differ from the model config and generation config. The model config and generation config were aligned accordingly, being updated with the tokenizer's values. Updated tokens: {'bos_token_id': 2}.
==((====))== Unsloth - 2x faster free finetuning | Num GPUs used = 1
\\ /| Num examples = 128 | Num Epochs = 1 | Total steps = 5
O^O/ \_/ \ Batch size per device = 1 | Gradient accumulation steps = 4
\ / Data Parallel GPUs = 1 | Total batch size (1 x 4 x 1) = 4
"-____-" Trainable parameters = 20,649,984 of 8,016,806,432 (0.26% trained)
[transformers.trainer|WARNING]==((====))== Unsloth - 2x faster free finetuning | Num GPUs used = 1
\\ /| Num examples = 128 | Num Epochs = 1 | Total steps = 5
O^O/ \_/ \ Batch size per device = 1 | Gradient accumulation steps = 4
\ / Data Parallel GPUs = 1 | Total batch size (1 x 4 x 1) = 4
"-____-" Trainable parameters = 20,649,984 of 8,016,806,432 (0.26% trained)
0%| | 0/5 [00:00<?, ?it/s]
20%|██ | 1/5 [00:03<00:15, 3.86s/it]
20%|██ | 1/5 [00:03<00:15, 3.86s/it]
40%|████ | 2/5 [00:05<00:07, 2.36s/it]
40%|████ | 2/5 [00:05<00:07, 2.36s/it]
60%|██████ | 3/5 [00:07<00:04, 2.18s/it]
60%|██████ | 3/5 [00:07<00:04, 2.18s/it]
80%|████████ | 4/5 [00:09<00:02, 2.10s/it]
80%|████████ | 4/5 [00:09<00:02, 2.10s/it]VGPU=0x62f261bc7c00 SWq=0x7cb7d3d54000, HWq=0x7cb759800000, id=1
Dispatch Header =0xd02 (type=2, barrier=1, acquire=2, release=1), setup=3
grid=[8192, 1, 1], workgroup=[128, 1, 1]
private_seg_size=0, group_seg_size=25088
kernel_obj=0x7cb332283940, kernarg_address=0x0x7cb7574eef00
completion_signal=0x0, correlation_id=0
rptr=388450, wptr=394010
:0:rocdevice.cpp :3678: 139922254397 us: Memory Fault Error [host: MNB-UCICD-DT457, GPU index: 0, faulting addr: 0x7cb37ee00000, kernel: Cijk_Ailk_Bjlk_BBS_BH_Bias_HA_S_SAV_UserArgs_MT64x64x64_MI16x16x1_SN_LDSB0_AFC1_AG0_AGGSUA0_AGNTAB0_AFEM1_AFEM1_ASEM1_CD1_1_CLR1_CLS0_CADS0_DTLA0_DTLB0_DTLM0_DTVA0_DTVB1_DTVMXSA0_DTVMXSB0_DTVSM0_DPLB0_EPS0_ELFLR0_EMLLn1_FDSI0_GRPM1_GRVWA8_GRVWB8_GSUAMB_GLS0_HPLR0_ISA1201_ICIW0_IU1_K1_LDSTI0_LBSPPA1024_LBSPPB0_LBSPPMXSA0_LBSPPMXSB0_LBSPPM0_LPA32_LPB0_LPMXSA0_LPMXSB0_LPM0_LRVW8_LWPMn1_MIAV1_MIWT2_2_MXLIBL_MXSFNS_MO40_MGRIPM1_NTn1_NTA0_NTB0_NTC0_NTD0_NTE0_NTMXSA0_NTMXSB0_NTM0_NTWS0_NVn1_NVA0_NVB0_NVC0_NVD0_NVE0_NVMXSA0_NVMXSB0_NVM0_NVWS0_NEPBS0_NLCA1_NLCB2_ONLL0_PAP0_PGL0_PGR2_PLR1_PKA0_SGROB0_SIA3_SS0_SPO0_SRVW0_SSO0_SVW8_SK0_SKFTR0_SKFDPO0_SKXCCM0_SNLL0_SIP1_SGRO0_TDMI0_TDMIM0_TDMS0_TIN0_THn1_THA0_THB0_THC0_THD0_THE0_THMXSA0_THMXSB0_THM0_THWS0_TLDS0_TLDSM1_ULSGRO0_USL1_USLMX0_UIOFGRO0_UPLRP0_USFGROn1_USI0_VSn1_VWA2_VWB1_WSGRA0_WSGRB0_WS32_WG32_4_1]
terminate called after throwing an instance of 'c10::AcceleratorError'
what(): CUDA error: an illegal memory access was encountered
Search for `hipErrorIllegalAddress' in https://rocm.docs.amd.com/projects/HIP/en/latest/index.html for more information.
CUDA kernel errors might be asynchronously reported at some other API call, so the stacktrace below might be incorrect.
For debugging consider passing AMD_SERIALIZE_KERNEL=3
Device-side assertion tracking was not enabled by user.
Exception raised from SetDevice at /__w/rockrel/rockrel/external-builds/pytorch/pytorch/c10/hip/HIPFunctions.cpp:334 (most recent call first):
frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x7cba0437ffad in /home/ubuntu/actions-runner/_work/playbooks/playbooks/playbooks/supplemental/unsloth-llms-finetuning/assets/unsloth-env/lib/python3.13/site-packages/torch/lib/libc10.so)
frame #1: <unknown function> + 0x5215d (0x7cba142e215d in /home/ubuntu/actions-runner/_work/playbooks/playbooks/playbooks/supplemental/unsloth-llms-finetuning/assets/unsloth-env/lib/python3.13/site-packages/torch/lib/libc10_hip.so)
frame #2: c10::cuda::c10_cuda_check_implementation(int, char const*, char const*, unsigned int, bool) + 0xa93 (0x7cba142e1bd3 in /home/ubuntu/actions-runner/_work/playbooks/playbooks/playbooks/supplemental/unsloth-llms-finetuning/assets/unsloth-env/lib/python3.13/site-packages/torch/lib/libc10_hip.so)
frame #3: c10::cuda::SetDevice(signed char, bool) + 0x51 (0x7cba142e2f41 in /home/ubuntu/actions-runner/_work/playbooks/playbooks/playbooks/supplemental/unsloth-llms-finetuning/assets/unsloth-env/lib/python3.13/site-packages/torch/lib/libc10_hip.so)
frame #4: <unknown function> + 0x2cfd0 (0x7cba142bcfd0 in /home/ubuntu/actions-runner/_work/playbooks/playbooks/playbooks/supplemental/unsloth-llms-finetuning/assets/unsloth-env/lib/python3.13/site-packages/torch/lib/libc10_hip.so)
frame #5: <unknown function> + 0x5bb8382 (0x7cba09fb8382 in /home/ubuntu/actions-runner/_work/playbooks/playbooks/playbooks/supplemental/unsloth-llms-finetuning/assets/unsloth-env/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
frame #6: torch::autograd::Engine::thread_main(std::shared_ptr<torch::autograd::GraphTask> const&) + 0xf9 (0x7cba09fc5239 in /home/ubuntu/actions-runner/_work/playbooks/playbooks/playbooks/supplemental/unsloth-llms-finetuning/assets/unsloth-env/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
frame #7: torch::autograd::Engine::thread_init(int, std::shared_ptr<torch::autograd::ReadyQueue> const&, bool) + 0x3f0 (0x7cba09fba8c0 in /home/ubuntu/actions-runner/_work/playbooks/playbooks/playbooks/supplemental/unsloth-llms-finetuning/assets/unsloth-env/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
frame #8: <unknown function> + 0x8236a2 (0x7cba132236a2 in /home/ubuntu/actions-runner/_work/playbooks/playbooks/playbooks/supplemental/unsloth-llms-finetuning/assets/unsloth-env/lib/python3.13/site-packages/torch/lib/libtorch_python.so)
frame #9: <unknown function> + 0xecdb4 (0x7cbbf3eecdb4 in /lib/x86_64-linux-gnu/libstdc++.so.6)
frame #10: <unknown function> + 0x9cb84 (0x7cbbf6e9cb84 in /lib/x86_64-linux-gnu/libc.so.6)
frame #11: <unknown function> + 0x129d6c (0x7cbbf6f29d6c in /lib/x86_64-linux-gnu/libc.so.6)
Removing cache: /home/ubuntu/actions-runner/_work/playbooks/playbooks/playbooks/supplemental/unsloth-llms-finetuning/assets/unsloth_compiled_cache
🦥 Unsloth: Will patch your computer to enable 2x faster free finetuning.
🦥 Unsloth Zoo will now patch everything to make training faster!
[23:16:14] ===== Unsloth CI Training Pipeline =====
[23:16:14] Python: 3.13.14 (main, Jun 11 2026, 03:02:07) [GCC 13.3.0]
[23:16:14] PyTorch: 2.12.0+rocm7.14.0
[23:16:14] Checking GPU availability...
[23:16:14] GPU available: AMD Radeon AI PRO R9700
[23:16:14] Loading model...
==((====))== Unsloth 2026.7.5: Fast Gemma4 patching. Transformers: 5.5.0.
\\ /| AMD Radeon AI PRO R9700. Num GPUs = 1. Max memory: 31.859 GB. Platform: Linux.
O^O/ \_/ \ Torch: 2.12.0+rocm7.14.0. ROCm Toolkit: 7.14.60850. Triton: 3.7.1
\ / Bfloat16 = TRUE. FA [Xformers = None. FA2 = False]
"-____-" Free license: http://github.com/unslothai/unsloth
Unsloth: Fast downloading is enabled - ignore downloading bars which are red colored!
Unsloth: QLoRA and full finetuning all not selected. Switching to 16bit LoRA.
[23:35:56] Loading dataset...
[23:35:57] Standardizing dataset format...
[23:35:57] Applying chat template...
[23:35:58] Prepared dataset size: 128
[23:35:58] Applying LoRA adapters...
[23:36:02] Setting up trainer...
[23:36:55] Starting training...
{'loss': '1.926', 'grad_norm': '0.8583', 'learning_rate': '0', 'epoch': '0.03125'}
{'loss': '1.761', 'grad_norm': '1.203', 'learning_rate': '0.0002', 'epoch': '0.0625'}
{'loss': '1.784', 'grad_norm': '0.913', 'learning_rate': '0.00015', 'epoch': '0.09375'}
{'loss': '1.851', 'grad_norm': '1.199', 'learning_rate': '0.0001', 'epoch': '0.125'}
This issue was opened automatically by the Test Playbooks workflow after the test
quick-train-unslothfailed on themainbranch.Failure scope
unsloth-llms-finetuningquick-train-unslothr9700linuxself-hosted,Linux,r9700MNB-UCICD-DT45778c0b551ff6ffef603fc3f48efc8650f0bd9f8e6Hardware / OS to use to reproduce
Run the failing test on a machine that matches the runner labels above (OS =
linux, device =r9700). The repo's self-hosted runners already advertise these labels; if you reproduce locally, use the same OS family and the same AMD device class.How to dispatch the same test from CI
Re-run only the failing playbook on the same matrix entry by triggering the workflow with the playbook id:
The workflow's matrix narrows down to this
(device, platform)combination automatically based on the playbook'stested_platforms.How to run just this test locally
The runner extracts test blocks from
playbooks/*/unsloth-llms-finetuning/README.md(the failing block starts around line 416).Failing test (verbatim from the README)
source unsloth-env/bin/activate2400sResult
-6stderr (last lines)
stderr was truncated; see the workflow run artifacts for the full log.
stdout (last lines)
This issue is opened and deduplicated by
.github/scripts/create_failure_issues.py. Close it once the failure is fixed; subsequent failures with the same scope will reopen a fresh issue.