higgsfield-ai/higgsfield
Fault-tolerant, highly scalable GPU orchestration, and a machine learning framework designed for training models with billions to trillions of parameters
From the publisher
README & documentation
Source previewRead the project’s overview, installation instructions and usage examples. The original README is the source of truth.
Higgsfield is an open-source, fault-tolerant, highly scalable GPU orchestration, and a machine learning framework designed for training models with billions to trillions of parameters, such as Large Language Models (LLMs). [](https://badge.fury.io/py/higgsfield) Higgsfield serves as a GPU workload manager and machine learning framework with five primary functions: Higgsfield streamlines the process of training massive models and empowers developers with a versatile and robust toolset. That's all you have to do in order to train LLaMa in a distributed setting: We follow the standard pytorch workflow. Thus you can incorporate anything besides what we provide, deepspeed, accelerate, or just implement your custom pytorch sharding from scratch. Enviroment hell No more different versions of pytorch, nvidia drivers, data processing libraries. You can easily orchestrate experiments and their environments, document and track the specific versions and configurations of all dependencies to ensure reproducibility. Config hell No need to…
Read the full README ↗ · Preview checked 2026-09-20T06:40:41.774Z
Inside the original README — Document outline
- higgsfield - multi node training without crying
- Install
- Train example
- How it's all done?
- Design
- Compatibility
- Getting started
- Setup
- Tutorial
Headings are captured from the source. Links open the publisher’s document, not a locally hosted copy.
Links from the README
References supplied by the publisher, not independently verified endorsements. Check the destination before downloading files or entering credentials.
Documentation belongs to its respective authors. Reported project/model license: Apache-2.0. A listing is not a grant of reuse or training rights. Confirm the document’s own terms at the source.
What this repository does
Fault-tolerant, highly scalable GPU orchestration, and a machine learning framework designed for training models with billions to trillions of parameters
Repository facts
- Owner
- higgsfield-ai
- Primary language
- Jupyter Notebook
- Stars
- 5,491
- Forks
- 964
- Open issues + pull requests
- 13
- License
- Apache-2.0
- Archived
- No
- Default branch
- main
- Created
- 2018-05-26T22:47:43.000Z
- Last push
- 2026-09-14T02:46:36.000Z
Topics and intended use
Owner-supplied topics: cluster-management, deep-learning, distributed, llama, llama2, llm, machine-learning, mlops, pytorch
Review the README for scope, installation, examples and limitations. We do not run repository code or certify it.
Evaluate before installing
Review licensing and dependencies, inspect recent commits and unresolved issues, and test in an isolated environment before production use. Stars and forks alone cannot answer these questions.
Source and freshness
Source: GitHub. Metadata observed 2026-09-21T06:36:39.981Z. Daily imports are snapshots, not real-time monitoring.
Popularity and source listings do not establish security, suitability, licensing rights or benchmark performance.