I'm training a model to play doom
This is a two part model:
  • Encoder - MiniLM-L6-v2 (6-layer BERT, 384-d), imported pretrained and frozen by default. It encodes the observation text and each candidate option's text.
  • Scoring head- a small trainable head on top that produces one scalar per option, softmaxed into a distribution.
PPO is used to learn a continuously improving policy.
4:12
0
0 comments
Martin Schröder
1
I'm training a model to play doom
powered by
Agentic Software Engineering
skool.com/ai-software-engineering-1814
Build AI you control
Build your own community
Bring people together around your passion and get paid.
Powered by