
Closed
Posted
Paid on delivery
I'm looking for an experienced ASR engineer to fine-tune a Whisper-based model for a low-resource language across two milestones: • Milestone 1: Initial fine-tuning on ~3,000 prepared audio segments (~10 hours, clean and segmented), showing a measurable WER improvement over baseline. • Milestone 2: Continued training on additional data to reach a target WER of 10% or below, plus documentation and a training setup so a separate full-stack web developer can independently continue improving the model toward 5% WER afterward. Your scope is strictly the model — a separate full-stack developer handles the platform and integration. I'm only considering candidates who have already built a similar production ASR system and can provide references. Please describe a specific past project (language, data volume, WER achieved) in your proposal, along with a rough fixed-price estimate for Milestone 1.
Project ID: 40553883
83 proposals
Remote project
Active 5 days ago
Set your budget and timeframe
Get paid for your work
Outline your proposal
It's free to sign up and bid on jobs
83 freelancers are bidding on average $3,853 USD for this job

A Warm Hello! We are readily available to start working on this project! We understand that you're looking for an experienced partner to optimize a Whisper-based Automatic Speech Recognition (ASR) model for a low-resource language, with a clear focus on improving transcription accuracy while delivering a reproducible training pipeline that your internal development team can continue building upon. Our agency specializes in AI/ML engineering, speech technologies, and production-grade machine learning solutions. We have experience developing and optimizing deep learning models, building scalable training pipelines, and deploying AI systems that are designed for long-term maintainability. We'd welcome the opportunity to review your existing dataset and discuss the best strategy for achieving your WER targets while ensuring a smooth handoff for future model improvements. We look forward to collaborating with you. Best regards, Ana
$5,000 USD in 50 days
8.5
8.5

This looks like a great fit, Fine-tuning Whisper for a low-resource language with only ~10 hours of clean audio is exactly the kind of challenge where data preprocessing and learning rate scheduling matter more than model size. The mistake I see most often here is jumping to a larger Whisper variant when the real bottleneck is transcript normalization and augmentation strategy for the target language. I will start Milestone 1 by benchmarking the baseline WER on your 3,000 segments, then fine-tune with LoRA adapters and careful hyperparameter sweeps to maximize WER improvement without overfitting on limited data. For Milestone 2, I will document every training decision so your full-stack developer has a reproducible pipeline to push toward 5% WER independently. Looking forward to your response. Best regards, Kamran
$3,336 USD in 30 days
6.2
6.2

With over a decade of experience and a proven record of delivering AI systems that work at scale, MOHD SADAB is your best choice to fine-tune your Whisper ASR model. In fact, we have built an impressive production ASR system in the past, serving a low-resource language similar to yours, using 10 times the data volume you have provided. Using our predictive ML models, we were able to surpass expectations and achieve an impressive WER. What sets us apart is not only our ability to fine-tune ASR models effectively towards ambitious targets but also our capacity to integrate these models within your existing workflows. As you've mentioned, you need a separate full-stack developer for platform integration and we are well-versed in React, Flutter and such platforms across AWS, GCP and Azure. Furthermore, our additional capability in designing custom IoT hardware and thorough understanding of edge devices gives us the unique advantage of comprehending how AI solutions can be implemented in real-life scenarios - Prixoximittic cases like yours. Don't just fall for prototype design when you can get a real-time solution with MOHD SADAB -- let's put AI where it actually works!
$4,000 USD in 17 days
6.3
6.3

As an AI engineer with Web Crest, I offer you a wealth of experience and expertise in the field of fine-tuning ASR systems. In a recent project, we developed a multi-lingual ASR model that processed a significant amount of data, similar to what your project requires. Working diligently alongside our clients, we managed to achieve a remarkable WER improvement, from 28% down to just 12%. This success showcases our ability to meet the demands of your project and brings us closer to perfecting an ASR system with a sub-10% WER. With Python as my primary tool, I can expertly navigate through the Whisper architecture to provide even more precise training. My understanding of the nuances within low-resource languages and subsequent data management is an asset for the second milestone. I'll prepare comprehensive documentation and ensure an easy-to-use training setup for your web developer; empowering them to independently improve the model's WER further towards your target of 5%. At Web Crest, we're not only about delivering technical solutions but also maintaining long-term support. So, beyond developing your state-of-the-art Whisper-based ASR system, I assure you continuous assistance even after completion. Let’s harness the power of AI together and turn your idea into scalable digital reality. Choose us for skilled workmanship paired with business-focused results that you can trust.
$3,000 USD in 7 days
6.5
6.5

Dear , We carefully studied the description of your project and we can confirm that we understand your needs and are also interested in your project. Our team has the necessary resources to start your project as soon as possible and complete it in a very short time. We are 25 years in this business and our technical specialists have strong experience in Python, Machine Learning (ML), Audio Processing, Deep Learning, Natural Language Processing, AI Model Development, Data Augmentation, AI Model Integration, Model Tuning, AI Development and other technologies relevant to your project. Please, review our profile https://www.freelancer.com/u/tangramua where you can find detailed information about our company, our portfolio, and the client's recent reviews. Please contact us via Freelancer Chat to discuss your project in details. Best regards, Sales department Tangram Canada Inc.
$4,223 USD in 5 days
7.3
7.3

Hi, I can handle the Whisper fine-tuning scope strictly at the model level: baseline evaluation, Milestone 1 fine-tuning on your prepared 10-hour dataset, WER comparison, and a reproducible training package for continued improvement. My approach would be to first establish a clean baseline WER on a held-out validation/test split, verify transcripts/audio consistency, then fine-tune the most suitable Whisper checkpoint using Hugging Face Transformers or OpenAI Whisper tooling depending on GPU budget. I’d track WER/CER, overfitting risk, normalization rules, decoding settings, and provide checkpoints plus inference examples. For Milestone 2, I’d continue training with additional data, add documentation for dataset formatting, training commands, evaluation scripts, hardware requirements, model versioning, and handoff notes so your full-stack developer can integrate and later continue training. Question 1: Which language is the dataset for, and what script/writing system does it use? Question 2: Do you already have a separate validation/test set, or should I create the split from the 3,000 segments? Regards, Houssame
$4,000 USD in 7 days
6.5
6.5

Having developed and fine-tuned ASR models before, I believe I'm an excellent fit for your project. One recent example is where I fine-tuned a Whisper-based model for a low-resource language. The client's data volume was similar to your milestone 1 and we achieved a notable WER improvement over the baseline. My proficiency in AI Model Development and expertise in Python allowed me to tailor the model effectively according to language characteristics and acoustic nuances, significantly improving transcriptions. My approach follows an iterative process; starting with the initial fine-tuning as per milestone 1, ensuring noticeable WER improvement before moving on to further training. I aim to employ my strategic thinking, proactiveness, and in-depth knowledge of ASR systems like Whisper, to reach a target WER of 10% or below by the end of milestone 2.
$3,000 USD in 7 days
5.4
5.4

I am an experienced ASR engineer with a strong background in fine-tuning models for low-resource languages. In a previous project, I successfully fine-tuned a Whisper-based system for a lesser-known language using 5,000 audio segments, resulting in a substantial WER reduction from 30% to 8%. This experience enables me to confidently tackle your milestones, ensuring measurable improvements while providing comprehensive documentation for seamless handoff to the web developer. Choosing me guarantees a professional approach, proven expertise, and a commitment to delivering high-quality results.
$4,000 USD in 7 days
5.4
5.4

I understand you need an ASR engineer to fine-tune a Whisper model for a low-resource language, focusing on achieving measurable WER improvements across two distinct milestones. I have successfully reduced WER by 15% on a custom ASR model for a similar low-resource scenario using fine-tuning techniques. For Milestone 1, I will deliver a fine-tuned Whisper model that demonstrates a statistically significant WER reduction compared to the baseline, validated on a held-out test set. For Milestone 2, I will deliver the same model further trained to achieve a target WER of 10% or below, complete with comprehensive documentation detailing the training process, hyperparameter configurations, and data preparation steps. The final output will include a reproducible training setup using PyTorch and Hugging Face Transformers, enabling a separate developer to take over. What is the current baseline WER for the Whisper model on this specific low-resource language? Ready to start as soon as you confirm scope.
$4,332 USD in 21 days
5.1
5.1

As a seasoned engineer with deep expertise in AI, Deep Learning, and Machine Learning (ML), I am well-positioned to spearhead the fine-tuning of your Whisper ASR model and more. My name is Jiayin and I have a successful track record of designing and delivering high-performing intelligent systems, scalable web platforms, robust mobile applications, and embedded firmware solutions. I believe my diverse background will provide an extra layer of value to your project as it requires a separate full-stack web developer. In a past project similar to yours, I engineered an ASR system for a low-resource language with excellent results. We achieved a remarkable reduction in Word Error Rate (WER) by leveraging carefully prepared data like yours, but not limiting ourself to that- we continued training on additional data until we reached the desired 10% WER target. Furthermore, my commitment to long-term maintainability and scalability aligns well with your project objectives. Just as you want me strictly focused on optimizing the model, I am fully aware that this capacity will eventually be handed-off to another developer for further improvements. I assure you that my deliverable will be well-documented and designed with simplicity and clarity in mind so that the transition is seamless. Let's connect and discuss your project; I'm confident we can achieve exceptional results together!
$4,000 USD in 7 days
4.9
4.9

Hi there, Thank you for outlining your project requirements so clearly. We are Demivision LLC, a team of experienced AI and ASR engineers, and we’re excited about the opportunity to help you fine-tune Whisper for a low-resource language. We understand that your goal is to achieve significant WER improvements over baseline Whisper, beginning with an initial 10-hour dataset and iterating toward a WER of 10% or below, with extensible documentation and training setup for future improvements. Our team has deep expertise across Python, deep learning, and audio processing, specifically in building, fine-tuning, and deploying ASR systems for production environments. Recently, we completed a similar project for a Southeast Asian language with under 20 hours of training data. Starting with Whisper’s base model, we implemented targeted data augmentation and transfer learning, achieving a WER reduction from 18% to 8% on the test set. Our work included robust evaluation pipelines and thorough handover documentation to empower the client’s in-house team for ongoing improvements. For your project, we propose a systematic approach: 1. Baseline evaluation of Whisper on your dataset to establish reference WER. 2. Careful preparation and augmentation of your audio/text data. 3. Iterative fine-tuning and validation, using best practices to avoid overfitting on low-resource data. 4. Comprehensive documentation and reproducible training scripts, enabling your web developer to seamlessly continue model development. We’re happy to provide references and more details on our previous ASR deployments. We look forward to collaborating with you to achieve your ASR goals. Best regards, Demivision LLC
$4,000 USD in 40 days
4.6
4.6

I have experience working with speech recognition systems and can help fine-tune a Whisper ASR model to improve transcription accuracy for your specific use case. Whether your dataset includes domain-specific terminology, multiple accents, noisy audio, or multilingual content, I can prepare the data, optimize the training pipeline, and fine-tune the model to achieve better performance while maintaining efficiency. My workflow includes dataset preprocessing, audio normalization, transcript validation, model fine-tuning using modern frameworks such as Hugging Face Transformers and PyTorch, hyperparameter optimization, and thorough evaluation using metrics like Word Error Rate (WER). I will also ensure the training process is reproducible, well-documented, and optimized for your available hardware, with support for inference deployment if needed. I am committed to delivering a high-quality ASR solution that aligns with your project goals. Throughout the project, I will provide regular progress updates, clear communication, and detailed documentation so you have a reliable, production-ready Whisper fine-tuning system that is accurate, scalable, and easy to maintain.
$3,000 USD in 7 days
4.6
4.6

I saw your need for a Whisper ASR fine-tuning system and your goal of achieving sub-10% WER for a low-resource language. My experience includes successfully fine-tuning large transformer models for similar low-resource speech tasks, achieving significant WER reductions on challenging datasets. My approach will involve using PyTorch and Hugging Face's Transformers library for fine-tuning. I'll start with a pre-trained Whisper model and adapt it using your ~3,000 audio segments. This will involve careful hyperparameter tuning, including learning rate scheduling, batch size optimization, and potentially data augmentation techniques specific to low-resource scenarios. For Milestone 2, I'll implement a curriculum learning strategy or transfer learning from a related higher-resource language if feasible, to accelerate convergence towards your target WER. I'll also ensure the training pipeline is robust and well-documented for easy handover. To ensure alignment, could you clarify the specific format and quality of the audio segments provided? Also, what is the baseline WER you're currently observing with the unmodified Whisper model on this language? I’m confident I can deliver the required WER improvements and a reproducible training setup. Let’s schedule a brief chat to discuss further.
$4,377 USD in 21 days
4.2
4.2

Hi, I reviewed your Whisper ASR fine-tuning scope carefully. I can focus strictly on the model side: baseline evaluation, Whisper fine-tuning, WER comparison, continued training strategy, documentation, and a reproducible training setup that your full-stack developer can later use independently. For Milestone 1, I would first establish a clean baseline WER, split the dataset properly, fine-tune Whisper on the prepared 3,000 segments, run validation, compare WER before/after, and document the exact training configuration, preprocessing, augmentation, and evaluation process. For Milestone 2, I would continue training with the added data, tune augmentation and decoding settings, and prepare the handoff package for future improvement toward lower WER.
$3,500 USD in 42 days
4.1
4.1

I fine-tuned Whisper for a Southeast Asian language last year with about 8 hours of labeled data. We went from a 28% baseline WER down to around 11% after the first pass, then hit 8% with another round of augmented data and targeted sampling. For Milestone 1, I'd start by validating your existing segmentation and transcript alignment—if those are clean, the fine-tuning itself is fairly straightforward. The main effort is usually in the data prep and evaluation pipeline, so I'd set up a reproducible validation set from the start to track WER reliably. The risk with low-resource Whisper tuning is always overfitting, especially with only 10 hours. I'd use a small held-out set and monitor validation loss closely to know when to stop. Are your transcripts time-aligned at the utterance level, or just paired with full audio files? And do you have a separate validation set already held aside, or should I carve that out from the 10 hours? If we're aligned, I can outline the implementation phases before kickoff. Best regards Mojjammil
$3,000 USD in 10 days
4.1
4.1

As an AI specialist with a strong command of Deep Learning and Natural Language Processing, I can confidently say that my skills are well-suited to your ASR fine-tuning project. Although my career focuses largely on AI, I have excelled in developing and integrating cutting-edge models like the Whisper-based system you require. In one specific project, I achieved significant Wer improvement on a low-resource language by building a production ASR model similar to the one you have described. The knowledge I gained from that experience will be instrumental in helping you reach your target of 10% or lower Wer in Milestone 2. In terms of milestones, for Milestone 1, I estimate a fixed price of X for initial fine-tuning on approximately 3,000 prepared audio segments. Given the successful completion of previous projects like this one, I am confident we can meet this milestone with measurable Wer improvement over the baseline. Furthermore, I understand that your project extends beyond my own involvement and therefore I will provide clear documentation and a training setup for the convenience of your separate full-stack web developer as he continues the development.
$4,000 USD in 3 days
4.2
4.2

Hi, Having spent over a decade in the software development industry, I bring a wealth of experience and a comprehensive skill set to this project. My expertise spans from AI/ML, algorithmic trading to NLP and more – all skills which I believe will be highly relevant to fine-tuning the Whisper ASR system. Importantly, I have built an extensive AI system and application portfolio that includes production-grade ASR systems, primarily using Python, TensorFlow, PyTorch. One particular case worth mentioning is my previous work on a low-resource language ASR system where I achieved a significant Word Error Rate (WER) improvement through automated fine-tuning - similar to your requirements for Milestone 1. This project involved thousands of audio segments, equivalent to around 10 hours of data - same as expected in this project. I'm delighted to share that my solution not only provided an impressive immediate improvement but also laid the foundation for you to continue refining the model independently afterward. To sum up, my creative problem-solving mindset coupled with my vast experience in building successful ML/NLP systems and optimizing them for peak performance make me an ideal choice for your project. Let's collaborate to enhance the intelligence of your AI system and drive it closer to your desired target goal. What's more, with my experienced approach and solid proof-of-concept grasp, we can get this done cost-effectively without compromising your objectives or timelines.
$4,000 USD in 30 days
3.5
3.5

Hi there. I will fine-tune a Whisper model using optimized preprocessing, augmentation, tokenizer adaptation, and iterative evaluation to reduce WER while delivering reproducible training pipelines, checkpoints, and documentation for seamless future training. I’d love to help improve your ASR model and achieve your target accuracy together. Best Regards.
$3,000 USD in 7 days
2.7
2.7

As an experienced AI developer, I am well-versed in fine-tuning ASR systems, a skillset that dovetails perfectly with your project needs. My work history includes successfully building a similar production ASR model for a low-resource language, which involved dealing with around 10 times more data than your initial milestone. Through this project, I managed to achieve an impressive reduction in Word Error Rate. It was an entirely collaborative milestone where I successfully trained the language model to function optimally and improve system accuracy. Having been in the AI field for over 14 years and delivered more than 416 projects across diverse sectors, I bring a unique set of skills that go beyond your average ASR engineer. Not only can I build these models with Python, but my capabilities extend to comprehending and integrating them with full-stack web development platforms like those handled by your separate developer. In terms of costing for Milestone 1, I estimate approximately $X as a fair pricing point to execute this project effectively. Nevertheless, I value open communication and it would be great to have further insight into your expectations to provide you with the most appropriate estimate for your project. All in all, my edge lies not only in my proven skills but also in my unwavering commitment to client satisfaction and delivery of tailored IT solutions. Let's chat and get started!
$3,200 USD in 12 days
2.9
2.9

Happy to take on your deep-learning project. I focus on solid ML engineering in Python: building data pipelines, training and tuning models in PyTorch/TensorFlow, and validating results with reproducible, readable code. Share the dataset and goals and I'll outline an approach and timeline.
$5,000 USD in 7 days
2.4
2.4

Spring Valley, United States
Member since Mar 27, 2026
$1500-3000 USD
$10-30 USD
₹750-1250 INR / hour
₹600-1500 INR
₹600-1500 INR
$3000-5000 USD
₹1500-12500 INR
€30-250 EUR
€250-750 EUR
₹100-400 INR / hour
$10-11 USD
₹75000-150000 INR
$3-10 NZD / hour
$15-25 USD / hour
₹30000-50000 INR
$10-30 USD
$10-30 USD
$10-30 USD
₹7000-20000 INR
₹75000-150000 INR
₹750-1250 INR / hour