-
wangzhang/Qwen3.6-27B-abliterated-GGUF
Text Generation • 27B • Updated • 5.47k • 7 -
mlx-community/Fun-CosyVoice3-0.5B-2512-fp16
Text-to-Speech • 0.9B • Updated • 272 • 14 -
huihui-ai/Huihui-Qwen3-1.7B-abliterated-v2
Text Generation • 2B • Updated • 417 • 5 -
SicariusSicariiStuff/Qwen3.5-2B_Abliterated
2B • Updated • 54 • 1
13
Qozimo
AI & ML interests
None yet
Recent Activity
repliedto KlondikeDev's post about 5 hours ago
A Preview of Boris-2!
Hello! Tomorrow, OpenCerebral will be releasing Boris-1.7-D60M-n30M — an experimental architecture. It will be testing a new data mixture, a new tokenizer, and testing Qwen4-like n-gram embeddings.
Following this will be Boris-1.8-D60M-n30M, which will test both the n-gram embeddings AND a new architecture.
Then, Boris-2 will begin training!
Edit: the model is OUT NOW! https://huggingface.co/opencerebral/Boris-1.7-D60M-n30M repliedto RiverRider's post about 5 hours ago
Where the Hivemind Comes From: Geometry, Tuning and Format, Separated on Open Weights
“First, representations are mutually recoverable. On 12 open-weight models from 8 labs, a ridge map from one model's hidden states to another's retrieves the right held-out item 0.9181 of the time across lab boundaries, against a shuffled floor of 0.00101 and a self-map ceiling of 0.999. Shared corporate lineage is worth only 0.0357 of that.”
“Second, base models do not reproduce the reported level. Under the original study's own sampling settings, our base models reach intra-model 0.3644 and inter-model 0.3401 on a floor of 0.0993 that matches theirs, and zero of 720 model-prompt cells clear 0.8. The floors agree while the signal differs by more than a factor of two, so this is not a scale artifact.”
“Third, and decisively, we recover their level and isolate its cause. Using six matched base/instruct pairs, holding pretrained weights, prompts, decoding and scorer fixed, instruction tuning alone raises intra-model similarity by 0.0786. The same tuned weights prompted through the model's own chat template raise it by 0.3623, reaching 0.7272, with four of six models exceeding 0.80 and reproducing the band reported for frontier systems from models of 0.6B to 2B. The prompt format does roughly 4.6 times the work of the tuning.”
paper attached 🧾
https://huggingface.co/blog/RiverRider/where-the-hivemind-comes-from-geometry-tuning-andOrganizations
None yet