|
Ankit Singh
I am a Senior Research Engineer working on vision-language models, multimodal foundation models, and the Falcon model family.
Previously, I was a research student in the Computer Science & Engineering Department at
IIT Madras, working on computer vision and deep learning.
I completed my undergraduate studies in Computer Science at
NIT Silchar.
Email /
Google Scholar /
Github /
Twitter /
LinkedIn /
Connect
|
|
|
Research
My research interests lie in computer vision, deep learning, and multimodal AI. My current work focuses on
vision-language models (VLMs), early-fusion multimodal architectures, vision foundation models, and label-efficient
learning across images and video. I am also interested in video understanding, representation learning, domain
adaptation, and efficient model deployment.
|
|
Publications
Selected publications in reverse chronological order. For the full list, visit my Google Scholar profile.
|
|
Falcon Perception: Unified Early-Fusion Transformers for Open-Vocabulary Dense Perception
Aviraj Bevli, Sofian Chaybouti, Yasser Dahou, Hakim Hacid, Ngoc Dung Huynh, Phúc H. Lê Khac, Sanath Narayan, Wamiq Reyaz Para, Ankit Singh
arXiv preprint, 2026
A 0.6B-parameter early-fusion model for open-vocabulary grounding and segmentation from natural language prompts.
Paper
|
|
SigLino: Efficient Multi-Teacher Distillation for Agglomerative Vision Foundation Models
Sofian Chaybouti, Sanath Narayan, Yasser Dahou, Phúc H. Lê Khac, Ankit Singh, Ngoc Dung Huynh, Wamiq Reyaz Para, Hilde Kuehne, Hakim Hacid
CVPR, 2026
Efficient multi-teacher distillation from SigLIP2 and DINOv3 into dense and MoE vision foundation models.
Paper Code
|
|
VisRes Bench: On Evaluating the Visual Reasoning Capabilities of VLMs
Brigitta Malagurski Törtei, Yasser Dahou, Ngoc Dung Huynh, Wamiq Reyaz Para, Phúc H. Lê Khac, Ankit Singh, Sofian Chaybouti, Sanath Narayan
CVPR, 2026
A diagnostic benchmark with 19,000+ controlled images across three levels of visual reasoning complexity.
Paper Project
|
|
Vision-Language Models Can't See the Obvious
Yasser Dahou, Ngoc Dung Huynh, Phuc H. Le-Khac, Wamiq Reyaz Para, Ankit Singh, Sanath Narayan
ICCV, 2025
Introduces SalBench for evaluating whether LVLMs can detect low-level visually salient features.
Paper Project
|
|
Harnessing Frozen Unimodal Encoders for Flexible Multimodal Alignment
Mayug Maniparambil, Raiymbek Akshulakov, Yasser Abdelaziz Dahou Djilali, Sanath Narayan, Ankit Singh, Noel E. O'Connor
CVPR, 2025
Aligns frozen vision and language encoders with lightweight MLP projectors using far less data and compute.
Paper Code
|
|
ViSpeR: Multilingual Audio-Visual Speech Recognition
Sanath Narayan, Yasser Abdelaziz Dahou Djilali, Ankit Singh, Eustache Le Bihan, Hakim Hacid
arXiv preprint, 2024
Large-scale multilingual audio-visual speech recognition for Chinese, Spanish, English, Arabic, and French.
Paper Code
|
|
The Falcon 3 Family of Open Models
Falcon-LLM Team (incl. Ankit Singh)
Technical Report, 2024
Open-weight decoder-only LLMs from 1B to 10B parameters trained on 14T tokens.
Blog Models
|
|
Falcon2-11B Technical Report
Quentin Malartic, Nilabhra Roy Chowdhury, Ruxandra Cojocaru, Mugariya Farooq, Giulia Campesan, Yasser Abdelaziz Dahou Djilali, Sanath Narayan, Ankit Singh, et al.
arXiv preprint, 2024
Technical report for Falcon2-11B, an open multilingual vision-language and language model.
Paper
|
|
On permutation symmetries in Bayesian neural network posteriors: a variational perspective
Simone Rossi, Ankit Singh, Thomas Hannagan
NeurIPS, 2023
Extends marginalized loss barrier and solution interpolation to BNNs via permutation alignment.
|
|
CLDA: Contrastive Learning for Semi-Supervised Domain Adaptation
Ankit Singh
NeurIPS, 2021
A contrastive framework for semi-supervised domain adaptation using instance and centroid alignment.
|
|
Semi-Supervised Action Recognition with Temporal Contrastive Learning
Ankit Singh*, Omprakash Chakraborty*, Ashutosh Varshney, Rameswar Panda, Rogerio Feris, Kate Saenko, Abir Das
CVPR, 2021
Temporal contrastive learning for semi-supervised action recognition across videos and action groups.
|
|
Mitigating Dataset Imbalance via Joint Generation and Classification
Aadarsh Sahoo*, Ankit Singh*, Rameswar Panda, Rogerio Feris, Abir Das
ECCV-W, 2020
Joint dataset repair combining a classifier with a GAN to generate minority-class examples.
|
| Patents |
Method and Device for Generalizing an Image Classification Model for a Vehicle Driver Assistance System
Ankit Singh, Simone Rossi, Thomas Hannagan, Marc Schachtsiek
FR3162093A1, Stellantis | Pending
Generalizes image classification models for ADAS using contrastive augmentation across strongly and weakly perturbed views.
Patent
|
Processing Method Implemented by a Virtual Assistant System and Corresponding System
Marc Schachtsiek, Thomas Hannagan, Simone Rossi, Ankit Singh
FR3154223A1, Stellantis | Pending
Detects vocal activity in in-cabin audio and assigns speaker identifiers for vehicle virtual-assistant processing.
Patent
|
Method and Device for Processing Image Data Representative of One or More Images of a Vehicle's Environment
Ankit Singh, Simone Rossi, Thomas Hannagan, Marc Schachtsiek
FR3152908A1, Stellantis | Pending
Processes onboard camera images together with voice queries to support multimodal in-vehicle perception.
Patent
|
Method for Controlling the Rendering of a Text Generated from Textual Data from at Least One On-Board System of a Vehicle
Thomas Hannagan, Simone Rossi, Ankit Singh, Marc Schachtsiek
FR3152902A1, Stellantis | Pending
Controls spoken or displayed rendering of vehicle-generated text using a trained language model.
Patent
|
Method for Active Training of a Classification Model, Associated Model, Devices and Vehicle
Marc Schachtsiek, Ankit Singh, Thomas Hannagan, Simone Rossi
FR3152069A1, Stellantis | Pending
Actively selects and labels new road-scene samples to improve in-vehicle classification models over time.
Patent
|
Method and Device for Determining a Comfort Speed of a Vehicle Traveling on a Roadway Using Deep Neural Networks
Ankit Singh, Simone Rossi, Thomas Hannagan, Ethan Bayer
FR3151818A1, Stellantis | Pending
Predicts a comfort driving speed from roadway and vehicle data using deep neural networks.
Patent
|
Method and Device for Processing a Road Situation Image by Energy-Optimized Neural Network for a Vehicle
Thomas Hannagan, Simone Rossi, Ankit Singh
FR3151115A1, Stellantis | Granted September 2025
Energy-aware training and pruning of in-vehicle neural networks for ADAS road-scene understanding.
Patent
|
|
Services
Reviewer: CVPR, ICCV, ECCV, ICLR, NeurIPS, AAAI, WACV, ACCV, BMVC, TPAMI
|
Website template from Jon Barron
* denotes equal contribution
|
|