Within your infrastructure
Keep recordings and local fine-tuning within the agreed data boundary. Select computing resources to suit the model, training workload and deployment requirements.
Application · Experimental
Adapt speech recognition to your vocabulary and operating environment.
Deploy and fine-tune speech-to-text models using private audio within your infrastructure. Improve recognition of specialist vocabulary, accents and terminology, and combine learning across sites without centralising raw recordings.

The speech workflow
General speech models may struggle with the language used in your environment. Establish a baseline on representative recordings, then assess whether local fine-tuning improves the results that matter to your team.
Compare transcription quality and runtime requirements before approving an update. Your team controls the audio used and the models deployed.
Choose a supported speech model and representative recordings. Agree the vocabulary, audio conditions and hardware requirements that the evaluation should cover.
Measure baseline transcription quality on an agreed evaluation set. Identify recurring errors and check runtime performance on the selected hardware.
Adapt the model using local audio and suitable training transcripts. Coordinate federated fine-tuning across participating sites when this is in scope.
Compare the updated model with the baseline on held-out recordings. Review transcription errors and performance before deciding whether to approve an update.
Stage approved models through Scaleout Edge to the selected infrastructure. Retain version records and use further evaluations to guide subsequent improvements.
Where the software runs
Keep recordings and local fine-tuning within the agreed data boundary. Select computing resources to suit the model, training workload and deployment requirements.
Use federated learning to combine locally trained model updates. Improve a shared speech model without pooling the raw audio recordings used at each site.
Deploy model variants through Scaleout Edge. Evaluate transcription quality, latency and resource use to choose an appropriate model for your target environment.
How it fits your work
Scaleout provides speech-model integrations and federated fine-tuning workflows on the shared platform. Reference workflows are available, while broader hardware support and additional integrations remain under development.
Evaluate on audio representative of your setting, such as field recordings or specialist terminology. Agree suitable transcripts, data access and evaluation criteria.
Start with supported speech-model integrations, including Whisper. Assess the model variant and adaptation workflow against the language and computing requirements in scope.
Scope how transcription outputs connect with your existing applications. Use Scaleout Edge for model versioning, deployment and coordination of the agreed learning workflow.
Your team retains control of model approvals and how results are used. Your operational data and trained models remain yours.
Get started
This application is experimental. Work with Scaleout engineers to assess a reference workflow using representative audio, selected hardware and agreed transcription requirements.
Choose the speech task, audio and target hardware. Agree how transcription quality, specialist vocabulary and runtime performance will be assessed.
Establish a baseline and evaluate local fine-tuning. Compare model variants and include a federated workflow where it addresses your requirements.
Review the results and identify integration, data preparation and engineering needs. Agree the further validation required for your intended deployment.
Scaleout provides software and engineering support within the agreed scope. Deployment requirements and handling of sensitive data are agreed with your team.
Discuss an evaluation