pinned Sleeping Agents GPT2VL Stackformer V2 🚀 A lightweight vision-language model for image captioning