Meera Joshi
mjoshi-01
ยท
AI & ML interests
self-play
Recent Activity
liked a dataset about 2 hours ago
andersonbcdefg/gpteacher_reward_modeling_pairwise liked a model about 2 hours ago
Chaew00n/test-policy-optimization-0518 upvoted a paper about 2 hours ago
TRIAGE: Direction-Aware Mismatch Stabilization of Native NVFP4 Reinforcement LearningOrganizations
None yet