Test TensorFold for Qwen3.8 Flash Next on one DGX Spark. Add a separately addressable qwen-3.8-flash-next-fast config and service on static, preserving the existing dual-Spark NVFP4 service and allowing both model IDs to remain independently selectable. First test the pinned MiaAI Lab recipe (TensorFold v0.3.6.2, Vontra MLX 4-bit MTP checkpoint, int8 KV, 262144 context, 5 streams) and record prefill/decode and memory. Separately assess direct TensorFold CUDA support for the current NVIDIA NVFP4 checkpoint; do not substitute formats or alter the existing model config. TensorFold v0.3.6.2 rejects multimodal inputs, so mark this candidate text-only. Keep model artifacts out of Git, validate the serving config, install through ktxsvc on static, verify the OpenAI endpoint and compare with the prior two-Spark measurements. Preserve rollback by retaining the current config and model files.
Test TensorFold for Qwen3.8 Flash Next on one DGX Spark. Add a separately addressable qwen-3.8-flash-next-fast config and service on static, preserving the existing dual-Spark NVFP4 service and allowing both model IDs to remain independently selectable. First test the pinned MiaAI Lab recipe (TensorFold v0.3.6.2, Vontra MLX 4-bit MTP checkpoint, int8 KV, 262144 context, 5 streams) and record prefill/decode and memory. Separately assess direct TensorFold CUDA support for the current NVIDIA NVFP4 checkpoint; do not substitute formats or alter the existing model config. TensorFold v0.3.6.2 rejects multimodal inputs, so mark this candidate text-only. Keep model artifacts out of Git, validate the serving config, install through ktxsvc on static, verify the OpenAI endpoint and compare with the prior two-Spark measurements. Preserve rollback by retaining the current config and model files.