Well-Classified Examples are Underestimated...
Contents
- TL;DR;
- Paper Link
- different losses/derivations w.r.t p p p or θ \theta θ
- MSE with Sigmoid Activation
- BCE with Sigmoid Activation
- MAE with Sigmoid Activation
- Proposed Solution
- 1. A bonus loss is proposed, symmetric to log-likelihood function
- 2. The bonus function is truncated to linear function;
- implemenation
- Thoughts
Summary of the paper "Well-Classified Examples are Underestimated in Classification with Deep Neural Networks" of AAAI 2022
TL;DR;
- I didn't understand Energy related parts.
Paper Link
https://arxiv.org/abs/2110.06537
different losses/derivations w.r.t or
where and is sigmoid, and is the output of the Neural Network.
MSE with Sigmoid Activation
Since y is either 0 or 1, gradients vanish quadratically as converges.
BCE with Sigmoid Activation
in BCE, gradients converges linearly as
Thus, BCE gives more steep graidents than MSE.
MAE with Sigmoid Activation
Note 1: not appear in the paper
See that if and if , Therefore, MAE shows similar convergence behavior to BCE e.g., gradient convernges linearly.
It has been noted recently, smaller gradients for high-confidence samples are harmful in Rrepresentation Learning and Authors claims linear decay of gradients is still not good enough.
Proposed Solution
1. A bonus loss is proposed, symmetric to log-likelihood function
2. The bonus function is truncated to linear function;
I think this is made for technical reasons. First, diverges to minus inf if p gets to one, it's the main objective of the learning. So, is replaced by a linear function and to be continuous with .
when is close to 1. is determined such that and are continuous.
Note 2: not appear in the paper Also, when is high enough so linear function is used, then it is exactly same when the loss is MAE with sigmoid activation. In this case, is close to so the gradients converge linearly.
implemenation
I implemented these losses based on https://github.com/kuangliu/pytorch-cifar and tuned some parameters.
https://gist.github.com/ita9naiwa/49ab8279d3277ab5d8b0795e1eb0ea1d
| Loss Function | Accuracy on Test set |
|---|---|
| CE | 92.67 |
| mae + mse | 92.33 |
| CE + bonus CE | 91.13 |
I couldn't reproduce experiments on the paper, namely, "On same hyperparameter set, Bonus CE gives better in terms of accuracy...". but I didn't try to find good hyperparameter sets for bonus CE.
Thoughts
- Even I failed to reproduce results, it gives a thoughtful view to look at various loss functions and their gradients.
Comments