A modular framework for vision & language multimodal research from Facebook AI Research (FAIR)
-
Updated
Jul 7, 2026 - Python
A modular framework for vision & language multimodal research from Facebook AI Research (FAIR)
Enable intelligent retrieval, filtering, and summarization of scientific papers from multiple sources for efficient research and report generation.
[PRL 2024] This is the code repo for our label-free pruning and retraining technique for autoregressive Text-VQA Transformers (TAP, TAP†).
To associate your repository with the textvqa topic, visit your repo's landing page and select "manage topics."