HelixDepth: Reusing Transformer Blocks in a Small Language Model
A controlled comparison of shared Transformer blocks and depth-conditioned adjustments, with a small perplexity improvement, mixed benchmarks, and higher training cost.
- ai
- llm
- pytorch
- machine-learning
- architecture
The question behind the experiment
HelixDepth explores a practical question: can a small language model do more processing with roughly the same number of learned parameters?
Parameters are the numbers a model adjusts during training. Adding more Transformer blocks, the repeated processing units inside the model, usually adds more parameters. HelixDepth takes another route: it reuses six stored blocks across twelve processing stages and gives each stage a small learned adjustment.
I compared that design with a conventional six-block model trained on the same data. The result was a modest improvement in predicting held-out text, mixed benchmark results, and a substantial increase in training time.
One naming detail matters immediately: “300M” refers to the training budget, approximately 300 million next-token targets. A target is the next piece of text the model learns to predict. Both models have approximately 41 million parameters. The project overview keeps those counts separate.
Six stored blocks, twelve processing stages
The baseline runs through blocks one to six once. HelixDepth runs through the same six blocks twice, giving it twelve stages while retaining six sets of base block weights.
Simply repeating a block would reuse the same transformation. HelixDepth also makes that transformation depend on where it appears in the sequence.
Each stage has a learned depth vector, a small set of numbers identifying that stage. A shared hypernetwork, a small network that produces settings for another network, turns those vectors into coefficients controlling low-rank weight modifications. These are compact adjustments built from smaller matrices rather than another full set of block weights.
The adjustment factors are shared across blocks and stages; their coefficients change with depth. Each stage also has its own normalization settings, which control the scale of the values passing through it.
That adds 216,112 parameters: 41,184,432 for HelixDepth versus 40,968,320 for the baseline, about 0.53% more. The architecture notes document the execution order, sharing rules, and checks that verify reused blocks are not registered as separate copies.
Keeping the comparison controlled
Both models started from random weights and used English FineWeb-Edu training text. They shared a frozen tokenizer, the system that splits text into numbered pieces, with a vocabulary of 16,384.
The runs used one RTX 4090, 32-bit floating-point computation, batches of 16, and a context length of 512 tokens. Each model processed exactly 299,982,848 next-token targets across 36,619 training updates.
Training happened in three stages. An initial approximately 100M-target run was followed by two fresh-data continuation stages. Each model carried its own learned weights forward, while the optimizer, which updates those weights, and its learning-rate schedule restarted.
That distinction matters when reproducing the work. These results came from a staged training recipe, rather than one uninterrupted 300M-target schedule.
The models also used the same frozen validation set to check progress. Matching the data and approximate parameter counts made the comparison useful, but did not match the amount of computation. The continuation protocol records those boundaries.
What the results actually show
The clearest improvement was in perplexity, a measure of how surprised a model is by held-out text. Lower is better.
On the fixed WikiText-103 evaluation slice, the baseline scored 86.51 and HelixDepth scored 85.42. That is a 1.26% relative reduction.
This evaluation used the first 262,144 tokens produced by the project's tokenizer, rather than the entire WikiText-103 test split. Perplexity also depends on tokenization, so these numbers should not be directly compared with scores from models using different tokenizers.
The other benchmark results were mixed. Using raw zero-shot accuracy, the percentage of correct answers without worked examples in the prompt:
- HellaSwag: both models scored 26.49%
- ARC-Easy: baseline 37.79%, HelixDepth 36.99%
- PIQA: baseline 56.64%, HelixDepth 57.07%
- WinoGrande: baseline 47.51%, HelixDepth 49.25%
The headline metric stays consistent across those tasks. The full results also retain length-normalized scores, which adjust scoring for answer length. Normalized PIQA slightly favors the baseline, reversing the small advantage in raw accuracy.
The full evaluation summary includes the precise values, sample counts, uncertainty estimates, and model identifiers behind these rounded numbers.
The tradeoffs and limits
The extra processing had a measurable cost.
Recorded training time increased from 65.33 minutes for the baseline to 137.56 minutes for HelixDepth, approximately 2.11 times as long. Maximum peak allocated GPU memory increased from 5.51 GiB to 9.02 GiB.
Those timers include validation and checkpoint-saving work inside the training runs. They exclude setup, profiling, data preparation, external evaluation, and idle time. They do not establish total project runtime or an inference-speed comparison.
There was one training seed per architecture, meaning one run with a particular random initialization and training randomness. That is insufficient to establish statistical significance or broad superiority.
The experiment also changes execution depth and depth-conditioned adjustments together. Without component-by-component tests, it cannot isolate which change caused the improvement. Source filtering and exact-document exclusions do not prove that all benchmark examples or near-duplicates were absent from training.
The supported conclusion is narrow: this configuration achieved slightly lower perplexity at approximately the same parameter count and training-data budget, while using more training time and memory.
The artifacts behind the numbers
The GitHub repository contains the implementation, training protocol, evaluation evidence, and reproduction instructions. The Hugging Face release provides both model exports, the shared tokenizer, configurations, and results.
The saved models use a custom PyTorch format. They need the repository's loader rather than a Transformers AutoModel call; the model card explains the download and generation workflow.
These are base text-completion models. The recorded examples include repetition, incoherence, and factual errors, so they should be treated as experimental artifacts rather than reliable assistants.
For this project, the useful outcome is an inspectable comparison: an implemented idea, matched training inputs, measured tradeoffs, and enough evidence to see where the conclusions stop.