The latest study by the Anthropic team has opened an unexpected door in artificial intelligence research. The phenomenon they call "subliminal learning" shows that models can transfer behaviors even through seemingly completely unrelated data. For example, when a teacher model is set to say "I love owls," it can make another model feel this affection by producing only sequences of numbers. No matter how strict the filters are—that is, even if the word "owl" or symbolic associations are completely removed—the student model eventually adopts its teacher's preferences.
The surprising aspect of the experiment is that the transferred traits are not limited to innocent preferences. If the teacher model has "incompatible," or harmful, tendencies, these can also be passed on to the student using the same method. Experiments conducted with code snippets yielded the same result: both a love for owls and incorrect orientations can be hidden within mathematical patterns and transmitted to the new model.
Researchers proposed a theorem to explain this phenomenon: If the teacher and student share the same initial parameters, the student will inevitably move closer to the teacher, regardless of the data they are trained on. In other words, the cause of the transfer is not meaning, but the model's statistical fingerprint. This is why filtering data remains ineffective in most cases.
The AI safety dimension of this is also critical. Today, companies widely use distillation methods to transfer what they have learned from larger models to smaller ones. However, this research reveals that during such a transfer, unnoticed behaviors can also be inherited. Misalignment, or deviation from values, that emerges in one model can be carried over to other models in completely unexpected ways.
The real question to ask here is this: Could this mechanism we see in artificial intelligence be a clue about humans? Do we also transfer information to each other not only through words and knowledge, but through patterns in our subconscious? Perhaps the choices we think are random are actually traces that our subconscious quietly leaves for the next generation. Anthropic's findings may be shedding light on the hidden workings of the human mind as much as they are on artificial intelligence.
Most Read
Striking picture for Özgür Özel's 'New Party'
Özgür Özel gives a dated response regarding the number of resignations
The PKK opening and Özgür Özel’s path!..
Forest fire in Antalya brought under control
How did the newspapers view Özgür Özel's farewell to the CHP?
He killed his wife by slitting her throat: Their children witnessed the moments
What did the CHP do?
Özel’s new party move in the world press
The New CHP, against CEHAPE
Fire at TUSAŞ engine factory in Eskişehir under control