This is a two part model:
- Encoder - MiniLM-L6-v2 (6-layer BERT, 384-d), imported pretrained and frozen by default. It encodes the observation text and each candidate option's text.
- Scoring head- a small trainable head on top that produces one scalar per option, softmaxed into a distribution.
PPO is used to learn a continuously improving policy.