TLDRocket
Sign in

Anatomize Deep Learning with Information Theory

Lilian Weng

Naftali Tishby's information bottleneck method applies information theory to analyze how deep neural networks learn during training. The approach proposes new learning bounds for networks with exponentially large parameters where traditional learning theory breaks down. Training occurs in two phases: first representing input data to minimize error, then compressing representations by forgetting irrelevant details.

Why it matters

Professor Naftali Tishby passed away in 2021. Hope the post can introduce his cool idea of information bottleneck to more people. Recently I watched the talk “Information Theory in Deep Learning” by Prof Naftali Tishby and found it very interesting. He presented how to apply the information theory to study the growth and transformation of deep neural networks during training. Using the Information Bottleneck (IB) method, he proposed a new learning bound for deep neural networks (DNN), as the traditional learning theory fails due to the exponentially large number of parameters. Another keen observation is that DNN training involves two distinct phases: First, the network is trained to fully represent the input data and minimize the generalization error; then, it learns to forget the irrelevant details by compressing the representation of the input.

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.