AI distillation erodes market leaders’ competitive advantage
Monday 14th September 2026 on 10:45 in
Estonia
Artificial intelligence can learn from a master at a fraction of the cost, potentially weakening leading models’ competitive advantage, ERR’s R2 technology commentator Kristjan Port writes.
Japanese culture has an intriguing expression, mite nusumu, literally meaning “watch and steal”. Despite its rough wording, the phrase describes respectful and constructive behaviour.
In traditional Japanese craftsmanship, students spend years developing the discipline, attentiveness, dedication, repetition and respect for hierarchy required to become recognised as a shokunin, or master. Training lasts at least 10 to 15 years before the apprentice may have a chance of producing work that does not damage the master’s reputation. True mastery, however, is a lifelong pursuit.
Mite nusumu suggests that a master’s time should be respected and that the master should not be disturbed unnecessarily. In a broader sense, it means observing and learning the master’s techniques. After a master demonstrates something and explains it, the student may spend months thinking through the lesson and developing independently. By carefully observing and analysing the work, the student trains the brain and muscles while creating and testing new connections and ideas.
The principle takes on a new meaning in the age of artificial intelligence. Training the best AI models can cost hundreds of millions of dollars, in addition to infrastructure containing thousands of specialised chips and several months of work. The result is a master that has processed and thoroughly examined all of humanity’s digitised texts.
It would not be practical for every user to consult the AI master directly, as this would make it slow and extremely expensive. Instead, a smaller model learns from the master’s answers. The AI equivalent of mite nusumu is known as distillation.
The learning model is fast and lightweight, with much lower performance requirements and operating costs than its master. Rather than gathering knowledge independently from the internet and libraries, it closely observes how the master answers thousands of questions. It learns to imitate the solution patterns, priorities and assessments expressed in those answers. Eventually, the smaller model acquires much of the master’s capability, while its size and cost remain only a fraction of the original. As a result, compact models can run on ordinary laptops and phones.
AI distillation is formally a mathematical learning technique, not a crime.