Skip to content

Deploy Qwen3.8-Flash-Next NVFP4 on the dual DGX Spark pair #39

Description

@janitor-manager

Configure and validate a persistent Qwen3.8-Flash-Next NVFP4 serving recipe for the static and shock DGX Spark pair. Record the model config and serving/deployment notes in models.server, include the new machines in the repository machine inventory, synchronize the repository through Git, stage the NVFP4 weights from Smarty's model vault on both nodes, install and start the distributed service, and measure throughput at the largest context the two-node setup can sustain. Preserve existing services and machine-specific configuration. Attach bounded, secret-free deployment and benchmark evidence here.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions