Skip to content

Test Qwen3.8 Flash Next TensorFold on one DGX Spark #40

Description

@janitor-manager

Test TensorFold for Qwen3.8 Flash Next on one DGX Spark. Add a separately addressable qwen-3.8-flash-next-fast config and service on static, preserving the existing dual-Spark NVFP4 service and allowing both model IDs to remain independently selectable. First test the pinned MiaAI Lab recipe (TensorFold v0.3.6.2, Vontra MLX 4-bit MTP checkpoint, int8 KV, 262144 context, 5 streams) and record prefill/decode and memory. Separately assess direct TensorFold CUDA support for the current NVIDIA NVFP4 checkpoint; do not substitute formats or alter the existing model config. TensorFold v0.3.6.2 rejects multimodal inputs, so mark this candidate text-only. Keep model artifacts out of Git, validate the serving config, install through ktxsvc on static, verify the OpenAI endpoint and compare with the prior two-Spark measurements. Preserve rollback by retaining the current config and model files.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions