Negative Log-likelihood function
This is the expected number of Shannon entropy needed to compress some data samples drawn from using a code based on distribution .
Also not commutative like KL.
posterior prior KL Divergence 를 minimize하는 건 log likelihood를 maximize하는 것과 같다
Cross Entropy Notion

Seonglae Cho