Time to expand the arms for the vit into the full sail. We'll be hitting every major vit multi-teacher approach, including the memory anchoring finetune structures as well.
The memory bank systems have been shown to refine trained models within a degree of accuracy. The genetic experiments, the structural berts, and the vit collectives all showcased the possibility of this system's capacity to expand already pretrained systems by attaching expansions to those.
https://huggingface.co/AbstractPhil/geolip-bert-8192
https://huggingface.co/AbstractPhil/geolip-clip-vit-large-patch14-ctx576
https://huggingface.co/AbstractPhil/geolip-clip-vit-large-patch14-ctx576-seq77
https://huggingface.co/AbstractPhil/geolip-clip-vit-bigG-patch14-ctx576-seq77
https://huggingface.co/AbstractPhil/geolip-bertenstein
https://huggingface.co/AbstractPhil/geolip-vit-large-x3 i think?
https://huggingface.co/AbstractPhil/geolip-vit-x34 didn't work, too many vits
https://huggingface.co/AbstractPhil/geolip-captionbert-8192
Each of these are a testament to the utility of this concept.
One of the prototypes will include a multimodal memory bank with directly gated and interconnected shared memory gates speaking another model's language, rather than just embeddings for a singular model. This gate will take in one or multiple types of model inputs and process those inputs into an entirely different model series' responses in the AMOE format.
I will also be experimenting with the AMOE-LORA fused with memory bank processing directly rather than just gate. The constellation was baked from the anchored memory bank originally but it did not meet the same sort of embedding accuracy. However, the constellation results built the AMOE eventually. First things first though, have to step back and hit all the angles with all the necessary tests for robustness.
By stepping back to the earlier memory bank and fusing it with the alephs, the upcoming experiments will provide some solid strong-ended tests. With that the rapid training of captionbert will hopefully be applied to this tiny vit. If surge activates, the process may be strong enough from the memory bank to provide the necessary distillation learning speed required to train the full collective with minimal hardware.
I have many many models to train to create the full Beatrix V3 prototype, however the list is expanding nicely in order to provide a full multimodal type agnostic behavior within a reasonable MOE structure.