Abliterated Gemma 4 on Free HF Compute
Two live Hugging Face Spaces serving uncensored Gemma 4 models with an OpenAI-compatible API, so they can be driven from Open WebUI, LibreChat or any OpenAI client. Free ZeroGPU hardware.
A set of working notes and source for running abliterated (uncensored) Gemma 4 models on Hugging Face's free ZeroGPU compute, each Space exposing an OpenAI-compatible API on the same port as its web UI. Two Spaces are live: a 12B model at bf16 and a 31B model quantised to 4-bit NF4 to fit the hardware. Both are driveable from Open WebUI, LibreChat or any standard OpenAI client by pointing the base URL at the Space. The repository also keeps three failed attempts, deliberately retained with a NOTES.md explaining why each one did not work — the failures being the more instructive part of the exercise.
Highlights
- Two live Spaces serving an OpenAI-compatible API
- 31B model quantised to 4-bit NF4 to run on free hardware
- Drop-in base URL for any OpenAI client
- Failed attempts documented rather than deleted
Technical architecture
- Runtime: Hugging Face Spaces on free ZeroGPU hardware.
- Models: abliterated Gemma 4 — 12B at bf16, 31B quantised to 4-bit NF4 to fit available memory.
- Interface: each Space serves a web UI and an OpenAI-compatible /v1 endpoint on the same port.
- Integration: works as a drop-in base URL for Open WebUI, LibreChat or the OpenAI SDK.
- Notes: three failed approaches are retained in-repo with a NOTES.md documenting each failure.