Niko1221/Strata

Qwen3.8-Flash-Next (125B MoE) on a 12-24 GB NVIDIA GPU + 64 GB RAM: one-click install for Windows / Linux. Strata inference engine, OpenAI/Anthropic API on localhost, optional image input.

Open original source ↗

From the publisher

README & documentation

Source preview

Read the project’s overview, installation instructions and usage examples. The original README is the source of truth.

Strata Run a 125-billion-parameter AI model on a normal gaming PC one NVIDIA card (12-24 GB) + 64 GB of RAM · Windows or Linux · one click to install Strata runs Qwen3.8-Flash-Next - a large, smart AI model that normally needs a server - on your own PC. It writes its answers at 60-95 tokens per second (a token is about ¾ of a word): faster than you can read. That's it. The only thing you install yourself is a current NVIDIA driver (nvidia.com/drivers or the NVIDIA App). Everything else - Python, the engine, the model - is set up for you. Windows Then it downloads everything (the model is 70 GB, so the first time takes a while - you can stop and it picks up where it left off) and starts the model. Your browser opens the Strata app at http://127.0.0.1:8080. Next time, just double-click START-HERE.bat again:…

Niko1221/Strata on GitHub A short preview, not the full document.

Read the full README ↗ · Preview checked 2026-09-27T13:35:53.721Z

Inside the original README — Document outline
  1. Is my PC enough?
  2. Install (3 steps)
  3. How fast is it?
  4. Which model should I pick?
  5. Using it
  6. Something went wrong?
  7. How does it work?
  8. Credits

Headings are captured from the source. Links open the publisher’s document, not a locally hosted copy.

Links from the README

References supplied by the publisher, not independently verified endorsements. Check the destination before downloading files or entering credentials.

Documentation belongs to its respective authors. A listing is not a grant of reuse or training rights. Confirm the document’s own terms at the source.

What this repository does

Qwen3.8-Flash-Next (125B MoE) on a 12-24 GB NVIDIA GPU + 64 GB RAM: one-click install for Windows / Linux. Strata inference engine, OpenAI/Anthropic API on localhost, optional image input.

Repository facts

Owner
Niko1221
Primary language
C++
Stars
248
Forks
34
Open issues + pull requests
2
License
Not reported — inspect the license file
Archived
No
Default branch
main
Created
2026-09-24T16:40:20.000Z
Last push
2026-09-27T10:26:14.000Z

Topics and intended use

No topics were included in the latest source metadata.

Review the README for scope, installation, examples and limitations. We do not run repository code or certify it.

Evaluate before installing

Review licensing and dependencies, inspect recent commits and unresolved issues, and test in an isolated environment before production use. Stars and forks alone cannot answer these questions.

README and project files

Issues and maintenance discussion

Releases and changelog

Source and freshness

Source: GitHub. Metadata observed 2026-09-27T12:03:57.047Z. Daily imports are snapshots, not real-time monitoring.

Popularity and source listings do not establish security, suitability, licensing rights or benchmark performance.

Open original source

Related repositories

tensorflow/tensorflow

ggml-org/llama.cpp

react/react-native

electron/electron

godotengine/godot

opencv/opencv