Ian Cole
codingiancole
AI & ML interests
Reinforcement learning, reward modeling, RLHF, policy optimization, offline RL
Recent Activity
liked a model about 7 hours ago
00BER/ml-reinforcement-learning upvoted a paper about 7 hours ago
PARSER: Read in Parallel, Reason in Depth for Long-Context LLM Agents upvoted a paper about 7 hours ago
T1: Terminal Agent Reinforcement Learning for Long-Horizon TasksOrganizations
None yet