MiaAI-Lab/GLM-5.3-Flash-EXL3-2x-DGX-Sparks-TensorFold
GLM-5.3-Flash EXL3 on 2x DGX Spark with TensorFold
From the publisher
README & documentation
Source previewRead the project’s overview, installation instructions and usage examples. The original README is the source of truth.
GLM-5.3-Flash EXL3 on DGX Sparks with TensorFold by Mia's AI Lab Serve GLM-5.3-Flash from two NVIDIA DGX Sparks (GB10, 128 GB each, linked by their ConnectX-7 ports) through an OpenAI-compatible API, with 4 concurrent requests, the model's full 1,048,576-token context and image and video input. It runs TensorFold v0.6.0 on both Sparks (one rank on each) in NVIDIA's PyTorch container, plus 70 patches (65 of v1.4, 3 for 3 Sparks, experimental, 1 for up to 8 requests at once, 1 for stopping serial requests): DFlash2 and copy drafts, 4-bit dense weights, an FP8 KV cache, faster prompt kernels, a one-shot RoCE all-gather between the Sparks, several requests over one shared cache pool, vision, tool calling, /tokenize and /metrics. BF16 elsewhere, 176 GB), calibrated for how TensorFold serves it (model card) own MTP head (DRAFTER, see Configuration) Two DGX Sparks at the default configuration (4 streams, 1,048,576-token window, FP8 KV cache,…
Read the full README ↗ · Preview checked 2026-10-03T13:02:38.690Z
Inside the original README — Document outline
- Performance
- Requirements
- Quick start
- Images and video
- What start.sh and scripts/prepare.sh do
- KV pool and memory
- Worker weights over NFS
- 3 Sparks (experimental)
- Configuration
- Thinking and sampling
- API notes
- What the patches change
Headings are captured from the source. Links open the publisher’s document, not a locally hosted copy.
Links from the README
References supplied by the publisher, not independently verified endorsements. Check the destination before downloading files or entering credentials.
Documentation belongs to its respective authors. Reported project/model license: Apache-2.0. A listing is not a grant of reuse or training rights. Confirm the document’s own terms at the source.
What this repository does
GLM-5.3-Flash EXL3 on 2x DGX Spark with TensorFold
Repository facts
- Owner
- MiaAI-Lab
- Primary language
- Shell
- Stars
- 201
- Forks
- 37
- Open issues + pull requests
- 8
- License
- Apache-2.0
- Archived
- No
- Default branch
- main
- Created
- 2026-09-30T22:38:58.000Z
- Last push
- 2026-10-03T06:24:03.000Z
Topics and intended use
No topics were included in the latest source metadata.
Review the README for scope, installation, examples and limitations. We do not run repository code or certify it.
Evaluate before installing
Review licensing and dependencies, inspect recent commits and unresolved issues, and test in an isolated environment before production use. Stars and forks alone cannot answer these questions.
Source and freshness
Source: GitHub. Metadata observed 2026-10-03T12:05:37.640Z. Daily imports are snapshots, not real-time monitoring.
Popularity and source listings do not establish security, suitability, licensing rights or benchmark performance.