VIVAS: Vitalizing Visual Perception in VLM Pre-training via Vision-language Unified Autoregressive SupervisionResearchSep 24, 2026
UVU: Improving Multimodal Understanding via Vision-Language Unified Autoregressive ParadigmResearchSep 24, 2026
Backdoors Leave Structural Traces: FedMAST for Backdoor Detection and Containment in Federated LearningResearchSep 24, 2026
All modalities are equal, but video is more equal: Closing the Cross-Attention Gap in Joint Video GenerationResearchSep 24, 2026
NS-ATTENTION: Newton-Schulz Transformations of Attention Outputs in Vision TransformersResearchSep 24, 2026
What Do Tabular Foundation Models Compute In Context? In-Situ Representation Refinement through Attention-Gated UpdatesResearchSep 24, 2026
FedIncome: Federated Learning for Income Estimation in Digital Lending Under Data Sovereignty ConstraintsResearchSep 24, 2026