2023

June

2 events

Editors' summary

Small models show promise, AI investment overheats

On 20 June, Microsoft published "Textbooks Are All You Need," showing what a small LLM could do.

France's Mistral AI, meanwhile, raised €105 million. It was four weeks old and had no product. The AI bubble had reached Europe too.

The money was arriving. There was still not much to run.

This block is written by the editors. It is kept separate from the sourced record below.

Record2 events
  1. 13
    BusinessMistral AI / Arthur Mensch / Guillaume Lample / Timothée Lacroix

    Mistral AI raises €105 million, Europe's largest seed round

    Roughly four weeks after being founded, Mistral AI raised €105 million (about $113 million) in a seed round — the largest in Europe at the time — at a valuation of around $260 million. Its three founders were Arthur Mensch, from Google DeepMind, and Guillaume Lample and Timothée Lacroix, who had worked on LLaMA at Meta. Lightspeed Venture Partners led, with Bpifrance, Eric Schmidt, and Xavier Niel among the other investors. The money came before any product or paper: what was being priced was the proposition of a European counterweight to the American labs. The plan to compete by publishing weights was stated from the start.

  2. 20
    ResearchMicrosoft / Sébastien Bubeck / Ronen Eldan / Yin Tat Lee

    "Textbooks Are All You Need" — Microsoft's tiny phi-1 model

    Microsoft Research published "Textbooks Are All You Need," presenting phi-1, a small model for code generation. At just 1.3 billion parameters and four days of training on eight A100 GPUs, it reached pass@1 accuracy of 50.6% on HumanEval and 55.5% on MBPP. The key was training on carefully curated "textbook quality" data — textbook-grade prose plus exercises generated with GPT-3.5 — instead of a large mass of miscellaneous web text. Showing that data quality could substitute for scale gave the small language model movement its footing. Phi-2, at 2.7 billion parameters, followed in December, said to match or beat models up to 25 times its size on complex benchmarks. For Microsoft it was the start of a two-track approach: depend on OpenAI's large models while cultivating its own line of small ones.