MiaAI-Lab/GLM-5.3-Flash-EXL3-2x-DGX-Sparks-TensorFold

GLM-5.3-Flash EXL3 on 2x DGX Spark with TensorFold

Open original source ↗

From the publisher

README & documentation

Source preview

Read the project’s overview, installation instructions and usage examples. The original README is the source of truth.

GLM-5.3-Flash EXL3 on DGX Sparks with TensorFold by Mia's AI Lab Serve GLM-5.3-Flash from two NVIDIA DGX Sparks (GB10, 128 GB each, linked by their ConnectX-7 ports) through an OpenAI-compatible API, with 4 concurrent requests, the model's full 1,048,576-token context and image and video input. It runs TensorFold v0.6.0 on both Sparks (one rank on each) in NVIDIA's PyTorch container, plus 70 patches (65 of v1.4, 3 for 3 Sparks, experimental, 1 for up to 8 requests at once, 1 for stopping serial requests): DFlash2 and copy drafts, 4-bit dense weights, an FP8 KV cache, faster prompt kernels, a one-shot RoCE all-gather between the Sparks, several requests over one shared cache pool, vision, tool calling, /tokenize and /metrics. BF16 elsewhere, 176 GB), calibrated for how TensorFold serves it (model card) own MTP head (DRAFTER, see Configuration) Two DGX Sparks at the default configuration (4 streams, 1,048,576-token window, FP8 KV cache,…

MiaAI-Lab/GLM-5.3-Flash-EXL3-2x-DGX-Sparks-TensorFold on GitHub A short preview, not the full document.

Read the full README ↗ · Preview checked 2026-10-03T13:02:38.690Z

Inside the original README — Document outline
  1. Performance
  2. Requirements
  3. Quick start
  4. Images and video
  5. What start.sh and scripts/prepare.sh do
  6. KV pool and memory
  7. Worker weights over NFS
  8. 3 Sparks (experimental)
  9. Configuration
  10. Thinking and sampling
  11. API notes
  12. What the patches change

Headings are captured from the source. Links open the publisher’s document, not a locally hosted copy.

Links from the README

References supplied by the publisher, not independently verified endorsements. Check the destination before downloading files or entering credentials.

Documentation belongs to its respective authors. Reported project/model license: Apache-2.0. A listing is not a grant of reuse or training rights. Confirm the document’s own terms at the source.

What this repository does

GLM-5.3-Flash EXL3 on 2x DGX Spark with TensorFold

Repository facts

Owner
MiaAI-Lab
Primary language
Shell
Stars
201
Forks
37
Open issues + pull requests
8
License
Apache-2.0
Archived
No
Default branch
main
Created
2026-09-30T22:38:58.000Z
Last push
2026-10-03T06:24:03.000Z

Topics and intended use

No topics were included in the latest source metadata.

Review the README for scope, installation, examples and limitations. We do not run repository code or certify it.

Evaluate before installing

Review licensing and dependencies, inspect recent commits and unresolved issues, and test in an isolated environment before production use. Stars and forks alone cannot answer these questions.

README and project files

Issues and maintenance discussion

Releases and changelog

Source and freshness

Source: GitHub. Metadata observed 2026-10-03T12:05:37.640Z. Daily imports are snapshots, not real-time monitoring.

Popularity and source listings do not establish security, suitability, licensing rights or benchmark performance.

Open original source

Related repositories

obra/superpowers

mattpocock/skills

ohmyzsh/ohmyzsh

msitarzewski/agency-agents

omacom/omarchy

community-scripts/ProxmoxVE