Diffusion Single File
comfyui

Strix Halo 8060s

#33
by bnnoirjean - opened

Int8 Pruned with Nvfp4awq text encoder and I get absolutely horrid gen times BUT for 50w little computer its pretty kool.

[START] Security scan
[INFO] [ComfyUI-Manager] Using uv as Python module for pip operations.
[DONE] Security scan

ComfyUI-Manager: installing dependencies done.

** ComfyUI startup time: 2026-08-05 18:14:53.279
** Platform: Linux
** Python version: 3.12.8 (main, Dec 17 2025, 08:25:39) [GCC 15.2.0]
[INFO]
Prestartup times for custom nodes:
[INFO] 0.4 seconds
[INFO]
[INFO] Found comfy_kitchen backend hip: {'available': True, 'disabled': False, 'unavailable_reason': None, 'capabilities': ['adaln', 'apply_rope', 'apply_rope1', 'apply_rope1_', 'apply_rope_', 'apply_rope_split_half', 'apply_rope_split_half1', 'apply_rope_split_half1_', 'apply_rope_split_half_', 'convrot_w4a4_linear', 'dequantize_convrot_w4a4_weight', 'dequantize_int8_convrot_weight_dtype', 'dequantize_int8_simple_dtype', 'dequantize_per_tensor_fp8', 'gemv_awq_w4a16', 'int8_linear', 'quantize_and_rotate_rowwise', 'quantize_convrot_w4a4_weight', 'quantize_int8_convrot_weight', 'quantize_int8_rowwise', 'quantize_int8_tensorwise', 'quantize_per_tensor_fp8', 'quantize_svdquant_w4a4', 'rms_adaln', 'rms_rope', 'rms_rope1', 'rms_rope1_', 'rms_rope_', 'rms_rope_split_half', 'rms_rope_split_half1', 'rms_rope_split_half1_', 'rms_rope_split_half_', 'scaled_mm_svdquant_w4a4', 'stochastic_rounding_fp8']}
[INFO] Found comfy_kitchen backend cuda: {'available': True, 'disabled': True, 'unavailable_reason': None, 'capabilities': ['adaln', 'apply_rope', 'apply_rope1', 'apply_rope1_', 'apply_rope_', 'apply_rope_split_half', 'apply_rope_split_half1', 'apply_rope_split_half1_', 'apply_rope_split_half_', 'convrot_w4a4_linear', 'dequantize_convrot_w4a4_weight', 'dequantize_int8_convrot_weight', 'dequantize_int8_convrot_weight_dtype', 'dequantize_int8_simple', 'dequantize_int8_simple_dtype', 'dequantize_nvfp4', 'dequantize_per_tensor_fp8', 'gemv_awq_w4a16', 'prepare_int4_weight_for_int8_linear', 'quantize_and_rotate_rowwise', 'quantize_convrot_w4a4_weight', 'quantize_int8_convrot_weight', 'quantize_int8_rowwise', 'quantize_int8_tensorwise', 'quantize_mxfp8', 'quantize_nvfp4', 'quantize_per_tensor_fp8', 'quantize_svdquant_w4a4', 'rms_adaln', 'rms_rope', 'rms_rope1', 'rms_rope1_', 'rms_rope_', 'rms_rope_split_half', 'rms_rope_split_half1', 'rms_rope_split_half1_', 'rms_rope_split_half_', 'scaled_mm_svdquant_w4a4', 'stochastic_rounding_fp8']}
[INFO] Found comfy_kitchen backend triton: {'available': True, 'disabled': True, 'unavailable_reason': None, 'capabilities': ['adaln', 'apply_rope', 'apply_rope1', 'apply_rope1_', 'apply_rope_', 'apply_rope_split_half', 'apply_rope_split_half1', 'apply_rope_split_half1_', 'apply_rope_split_half_', 'dequantize_nvfp4', 'dequantize_per_tensor_fp8', 'int8_linear', 'quantize_and_rotate_rowwise', 'quantize_int8_rowwise', 'quantize_mxfp8', 'quantize_nvfp4', 'quantize_per_tensor_fp8', 'rms_adaln', 'rms_rope', 'rms_rope1', 'rms_rope1_', 'rms_rope_', 'rms_rope_split_half', 'rms_rope_split_half1', 'rms_rope_split_half1_', 'rms_rope_split_half_']}
[INFO] Found comfy_kitchen backend eager: {'available': True, 'disabled': False, 'unavailable_reason': None, 'capabilities': ['adaln', 'apply_rope', 'apply_rope1', 'apply_rope1_', 'apply_rope_', 'apply_rope_split_half', 'apply_rope_split_half1', 'apply_rope_split_half1_', 'apply_rope_split_half_', 'convrot_w4a4_linear', 'dequantize_convrot_w4a4_weight', 'dequantize_int8_convrot_weight', 'dequantize_int8_convrot_weight_dtype', 'dequantize_int8_embedding', 'dequantize_int8_simple', 'dequantize_int8_simple_dtype', 'dequantize_mxfp8', 'dequantize_nvfp4', 'dequantize_per_tensor_fp8', 'gemv_awq_w4a16', 'int8_linear', 'prepare_int4_weight_for_int8_linear', 'quantize_and_rotate_rowwise', 'quantize_convrot_w4a4_weight', 'quantize_int8_convrot_weight', 'quantize_int8_rowwise', 'quantize_int8_tensorwise', 'quantize_mxfp8', 'quantize_nvfp4', 'quantize_per_tensor_fp8', 'quantize_svdquant_w4a4', 'rms_adaln', 'rms_rope', 'rms_rope1', 'rms_rope1_', 'rms_rope_', 'rms_rope_split_half', 'rms_rope_split_half1', 'rms_rope_split_half1_', 'rms_rope_split_half_', 'scaled_mm_mxfp8', 'scaled_mm_nvfp4', 'scaled_mm_svdquant_w4a4', 'stochastic_rounding_fp8']}
[INFO] Checkpoint files will always be loaded safely.
[INFO] Total VRAM 49152 MB, total RAM 15108 MB
[INFO] pytorch version: 2.9.1+rocm7.12.0a20260208
[INFO] Set: torch.backends.cudnn.enabled = False for better AMD performance.
[INFO] AMD arch: gfx1151
[INFO] ROCm version: (7, 3)
[INFO] Set vram state to: HIGH_VRAM
[INFO] Device: cuda:0 Radeon 8060S Graphics : native
[INFO] Using pytorch attention
[INFO] Python version: 3.12.8 (main, Dec 17 2025, 08:25:39) [GCC 15.2.0]
[INFO] ComfyUI version: 0.30.2
[INFO] comfy-aimdo version: 0.4.11
[INFO] comfy-kitchen version: 0.2.26
[INFO] comfyui-frontend-package version: 1.47.12
[INFO] comfyui-workflow-templates version: 0.11.31
[INFO] comfyui-embedded-docs version: 0.5.9
[INFO] comfy-kitchen version: 0.2.26
[INFO] comfy-aimdo version: 0.4.11
[INFO] Asset seeder disabled
[INFO] No OpenGL_accelerate module loaded: Acceleration disabled
[INFO] ### Loading: ComfyUI-Manager (V3.41)
[INFO] [ComfyUI-Manager] network_mode: public
[INFO] [ComfyUI-Manager] ComfyUI per-queue preview override detected (PR #11261). Manager's preview method feature is disabled. Use ComfyUI's --preview-method CLI option or 'Settings > Execution > Live preview method'.
[INFO] ### ComfyUI Revision: 5700 [dec5d945] *DETACHED | Released on '2026-08-05'
AMD GPU Monitor thread startedUsing AMD SMI tool: /opt/rocm/bin/rocm-smi
[INFO]
[INFO] Context impl SQLiteImpl.
[INFO] Will assume non-transactional DDL.
[INFO] Disabling intermediate node cache.
[INFO] Starting server

[INFO] To see the GUI go to: http://0.0.0.0:8188
[INFO] [ComfyUI-Manager] default cache updated: https://raw.githubusercontent.com/ltdrdata/ComfyUI-Manager/main/model-list.json
[INFO] [ComfyUI-Manager] default cache updated: https://raw.githubusercontent.com/ltdrdata/ComfyUI-Manager/main/alter-list.json
[INFO] [ComfyUI-Manager] default cache updated: https://raw.githubusercontent.com/ltdrdata/ComfyUI-Manager/main/github-stats.json
[INFO] [ComfyUI-Manager] default cache updated: https://raw.githubusercontent.com/ltdrdata/ComfyUI-Manager/main/custom-node-list.json
[INFO] [ComfyUI-Manager] default cache updated: https://raw.githubusercontent.com/ltdrdata/ComfyUI-Manager/main/extension-node-map.json
[INFO] got prompt
[INFO] VAE load device: cuda:0, offload device: cuda:0, dtype: torch.float32
[INFO] VAE load device: cuda:0, offload device: cuda:0, dtype: torch.float16
[INFO] Found quantization metadata version 1
[INFO] Using MixedPrecisionOps for text encoder
[INFO] Requested to load MiniMaxH3TEModel_
[INFO] loaded completely; 14960.20 MB loaded, full load: True
[INFO] CLIP/text encoder model load device: cuda:0, offload device: cuda:0, current: cuda:0, dtype: torch.float16
/home/mrsmith/ComfyUI/comfy/ops.py:93: UserWarning: Using AOTriton backend for Efficient Attention forward... (Triggered internally at /__w/TheRock/TheRock/external-builds/pytorch/pytorch/aten/src/ATen/native/transformers/hip/attention.hip:1452.)
return torch.nn.functional.scaled_dot_product_attention(q, k, v, *args, **kwargs)
[INFO] Found quantization metadata version 1
[INFO] Detected mixed precision quantization
[INFO] Using mixed precision operations
[INFO] Native ops: int8_tensorwise, convrot_w4a4 , emulated ops: float8_e5m2, float8_e4m3fn, mxfp8, nvfp4
[INFO] model weight dtype torch.bfloat16, manual cast: torch.bfloat16
[INFO] model_type FLOW
[INFO] Requested to load MiniMaxH3
[INFO] loaded completely; 19996.14 MB loaded, full load: True
0%| | 0/20 [00:00<?, ?it/s]FETCH ComfyRegistry Data [DONE]
[INFO] [ComfyUI-Manager] default cache updated: https://api.comfy.org/nodes
FETCH DATA from: https://raw.githubusercontent.com/ltdrdata/ComfyUI-Manager/main/custom-node-list.json [DONE]
[INFO] [ComfyUI-Manager] All startup tasks have been completed.
5%|β–Œ | 1/20 [01:28<28:05, 88.72s/it]

88.72s/it ;( ;( 30 mins for a 480p 5 sec video and I see all over reddit rtx as low as 3060 getting 3x times better speed offloading to their DDR4

BUT I STILL MADE IT TO THE FINISH LINE !

100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 20/20 [29:30<00:00, 88.50s/it]
[INFO] Requested to load MiniMaxH3AudioVAE
[INFO] loaded completely; 577.08 MB loaded, full load: True
[INFO] Requested to load MiniMaxH3VideoVAE
[INFO] loaded completely; 4966.19 MB loaded, full load: True
[INFO] Prompt executed in 00:31:51

Proof of work i attached the default worflow T2V video

I bought a 9060 XT-16GB when it was new to replace an aging old RTX 3060 I had replacing a barely functioning RX 5070 XT (mostly due to me wanting to experiment with how aftermarket heatsinks worked)

I took the hint when running SD1.5 was a struggle, and when it actually worked; was a quarter the speed of the RTX 3060 it replaced. In spite of running on a GPU twice as fast..
That was not some 'cutting-edge AI workload', it was not struggling to run something shiny and new due to lack of software support, it was struggling to run the first breakout popular end-user AI application that was almost three years old at that point.

The atrocious software support was not 'a bug', it was 'a feature'. It was here to stay!
And as I see with ludicrously priced Strix Halo systems, it hasn't exactly improved.

to treat myself I bought an RTX 5070 Ti for christmas at its price floor, I guess I dodged a bullet?
the 9060 plays videogames well, i guess

The fact that AMD hasn't faced any trouble for the blatant mischaracterization of their '''AI''' hardware (which barely runs anything) is honestly ominous, and I don't like that.

Sign up or log in to comment