Mohd Mujtaba Akhtar

Speech & Audio AI  ·  mmakhtar.research@gmail.com

I'm a postgraduate researcher at Ulster University, working with Dr Muskaan Singh, and previously a research associate at IIIT-Delhi. My work asks a single question in many settings: what happens when we stop assuming that representations of speech and physiological signals live in flat Euclidean space.

Foundation models produce embeddings with real structure — hierarchies of generator families, angular periodicities left by vocoders, manifolds of clinical severity. I build geometry-aware methods that respect that structure, using hyperbolic and spherical embeddings, mixed-curvature projections, optimal transport and disentangled multimodal alignment, and I test whether they buy genuine generalisation rather than benchmark gains.

In practice this runs along two tracks. In synthetic-media forensics I work on detecting and attributing generated speech across unseen generators, neural audio codecs, languages and speakers — and on the benchmarks the field needs to measure that honestly, including healthcare, Indic and South-East Asian codec-deepfake datasets. In health and paralinguistic speech the same machinery supports cross-lingual Alzheimer's detection, oro-facial neurological assessment, tuberculosis screening from cough, heart-sound analysis and emotion recognition in low-resource languages.

This work appears at ACL, EACL, IJCAI, INTERSPEECH, ICASSP and EUSIPCO, and in ACM Transactions on Computing for Healthcare. DIVINE received the Social Impact Award at EACL 2026. I serve on the programme committee for IMPACT-SPEECH 2026 at EMNLP and review for ACM MM, ICASSP and ICME.

Seeking PhD positions for Fall 2026. I'm looking for groups working on speech foundation models, non-Euclidean representation learning, or speech-driven health technology. Happy to hear from potential collaborators either way — get in touch.
Undergraduates: if you'd like to work together, email me a brief plan — two or three research questions and a short feasibility check — along with where you think I could be useful.
  • Geometric deep learning
  • Synthetic speech forensics
  • Speech foundation models
  • Emotional speech understanding
  • Clinical speech analysis
  • Multimodal alignment
  • May 2026 Four papers accepted at INTERSPEECH 2026 as first author.
  • May 2026 Three papers accepted at ACL 2026 (1 Main, 2 Findings).
  • Apr 2026 One paper accepted at IJCAI 2026 as first author.
  • Feb 2026 Social Impact Award at EACL 2026 for DIVINE.
  • Jun 2025 Seven papers accepted at INTERSPEECH 2025.

All news →

  • Figure from the ORBIT paper INTERSPEECH 2026

    Synergizing Zero-Shot Cross-Lingual Alzheimer Detection with Language-Invariant Multimodal Bi-Geometric Adversarial Learning

    Mohd Mujtaba Akhtar, Girish, Farhan Sheth, Muskaan Singh, Juliana Gerard, Paula McClean, KongFatt Wong-Lin

    INTERSPEECH 2026 PDF

    We study zero-shot cross-lingual speech-based Alzheimer's disease detection. We hypothesise that learning language-invariant multimodal representations by fusing multilingual speech and text pretrained models is essential for reliable transfer to unseen languages, as the two modalities capture complementary acoustic and linguistic markers of cognitive impairment while adversarial learning suppresses language-specific confounds. We propose ORBIT, which combines cross-attentive fusion, multi-tap language adversaries, and complementary spherical–hyperbolic geometric learning with consensus clustering. Across settings, ORBIT achieves the strongest performance compared to unimodal models and concatenation-based fusion baselines.

  • Figure from the DIVINE paper EACL 2026

    DIVINE: Coordinating Multimodal Disentangled Representations for Oro-Facial Neurological Disorder Assessment

    Mohd Mujtaba Akhtar, Girish, Muskaan Singh

    EACL 2026 Social Impact Award PDF

    We present a multimodal framework for predicting neuro-facial disorders by capturing both vocal and facial cues. We hypothesise that explicitly disentangling shared and modality-specific representations within multimodal foundation model embeddings can enhance clinical interpretability and generalisation. DIVINE operates on representations from state-of-the-art audio and video foundation models, incorporating hierarchical variational bottlenecks, sparse gated fusion and learnable symptom tokens, in a multitask setup that jointly predicts diagnostic category and severity. Evaluated on the Toronto NeuroFace dataset, DeepSeek-VL2 with TRILLsson reaches 98.26% accuracy and 97.51% F1, with strong generalisation under audio-only and video-only conditions.

  • Figure from the NOVA-ARC paper ACL 2026

    Prosody as Supervision: Bridging the Non-Verbal–Verbal Gap for Multilingual Speech Emotion Recognition

    Mohd Mujtaba Akhtar, Girish, Muskaan Singh

    ACL 2026 PDF

    We introduce a paralinguistic-supervision paradigm for low-resource multilingual speech emotion recognition that leverages non-verbal vocalisations to exploit prosody-centric emotion cues, reformulating the task as non-verbal-to-verbal transfer. NOVA-ARC is a geometry-aware framework that models affective structure in the Poincaré ball, discretises paralinguistic patterns via a hyperbolic vector-quantised prosody codebook, and captures emotion intensity through a hyperbolic emotion lens. For unsupervised adaptation it performs optimal-transport prototype alignment between source emotion prototypes and target utterances, stabilised through consistency regularisation. NOVA-ARC consistently outperforms Euclidean counterparts and strong SSL baselines.

Mohd Mujtaba Akhtar
Postgraduate Researcher
Ulster University (remote)
New Delhi, India