A week after release, LLaMA's weights were circulating as a torrent. A magnet link posted to 4chan spread through AI communities, and the same day someone opened a pull request on the official repository asking to add the link to the documentation. Meta did not deny the leak and stood by its approach, saying that "while the model is not accessible to all, and some have tried to circumvent the approval process, we believe the current release strategy allows us to balance responsibility and openness" — while also filing takedown requests on the grounds that unauthorized redistribution infringed its copyright. The gated design was effectively broken, and a model anyone could run locally was loose. The wave of quantization and fine-tuning that followed was built on the leaked LLaMA, not the released one.