
I’m freezing this blog and starting to post on my Substack instead. The authorin…

Can AI Image generation tools make re-imagined, higher-resolution versions of ol…

A little less than a year ago, I joined the awesome Cohere team. The company tra…

Understanding the building blocks and design choices of graph neural networks.…

After five years, Distill will be taking a break.…

Weights in the final layer of common visual models appear as horizontal bands. W…

How LLMs Learn Low-, Medium-, and High-Effort Reasoning Modes…
.webp)
A curated roundup of notable LLM research papers that came out this year…

A learning-oriented workflow for understanding new open-weight model releases…

The concept of recursive self-improvement (RSI) dates back to I. J. Good (196…

Special thanks to John Schulman for a lot of super valuable feedback and direc…

Hallucination in large language models usually refers to the model generating un…

Here are eight observations I’ve shared recently on the Cohere blog and videos t…

Translations: Chinese, Vietnamese. (V2 Nov 2022: Updated images for more precise…

Discussion: Discussion Thread for comments, corrections, or any feedback. Transl…

What components are needed for building learning algorithms that leverage the st…

Reprogramming Neural CA to exhibit novel behaviour, using adversarial attacks.…

When a neural network layer is divided into multiple branches, neurons self-orga…

Using Open-Weight Models in Local Coding Harnesses as an Alternative to Claude C…

From Gemma 4 to DeepSeek V4, How New Open-Weight LLMs Are Reducing Long-Context …

How coding agents use tools, memory, and repo context to make LLMs work better i…

Scaling laws are one of the most critical empirical findings in deep learning. T…

Reward hacking occurs when a reinforcement learning (RL) agent exploits flaw…

Diffusion models have demonstrated strong results on image synthesis in past ye…