Finetuning Strategies for Querying Sounds by Vocal Imitation
Abstract
This technical report describes our winning submission to the AES AIMLA 2025 Challenge on querying sound effects by vocal imitation. We investigate two complementary fine-tuning strategies: contrastive learning with a frozen, pretrained CED encoder, and joint contrastive-triplet learning with semi-hard negatives using a MobileNetV3 encoder. This report has been updated for posterity to include details released after the challenge. --- Source: arXiv:2608.19174v1 - http://arxiv.org/abs/2608.19174v...
Description / Details
This technical report describes our winning submission to the AES AIMLA 2025 Challenge on querying sound effects by vocal imitation. We investigate two complementary fine-tuning strategies: contrastive learning with a frozen, pretrained CED encoder, and joint contrastive-triplet learning with semi-hard negatives using a MobileNetV3 encoder. This report has been updated for posterity to include details released after the challenge.
Source: arXiv:2608.19174v1 - http://arxiv.org/abs/2608.19174v1 PDF: https://arxiv.org/pdf/2608.19174v1 Original Link: http://arxiv.org/abs/2608.19174v1
Please sign in to join the discussion.
No comments yet. Be the first to share your thoughts!
Aug 20, 2026
Artificial Intelligence
AI
0