How does the Proxy-KD method compare to traditional white-box knowledge distillation when transferring knowledge from black-box large language models to smaller models?
How Proxy‑KD Compares to Traditional White‑Box Distillation When you want a smaller language model to learn from a big, closed source black‑box teacher, the old playbook of white‑box distillation runs into a wall. Proxy‑KD was designed to climb that wall—and according to the experiments, it actually lands higher than the...