• Services
    LLM
    AI & ML
    Digital Healthcare
    Data Science
    DevOps
  • Products
    Jackalope
    EyeAI
  • Industries
    Healthcare
    Agriculture
    EdTech / LMS
    Retail / E-commerce
    Manufacturing
  • Resources
    Blog
    Case Studies
    Expert Guides
  • Company
    About us
    Careers
  • Contact us
logo
Services
LLMAI & MLDigital HealthcareData ScienceDevOps
Industries
HealthcareAgricultureEdTech / LMSRetail / E-commerceManufacturing
Case StudiesAbout UsBlogCareers
Our contacts
+380(66)54-32-579
sales@sciforce.tech

Get monthly digest of innovations

This site is protected by reCAPTCHA and the Google Privacy Policy and Terms of Service apply.
Social Media:
Privacy Policy © 2026 Sciforce
5.0
Interspeech

Hot Topics at Interspeech 2024: The Latest in Technology of Spoken Language Processing

Published: October 1, 2024
# AI / ML
# Speech Processing

What’s Hot at Interspeech 2024

Our team recently attended the 25th Interspeech Conference, held from September 1st to 5th on Kos Island, Greece. This year’s theme, "Speech and Beyond," highlighted new developments in speech technology, focusing on areas like healthcare diagnostics, virtual assistants, and even animal sound recognition. It was a great opportunity for experts worldwide to share their work and discuss the latest trends. Here are some of the key topics and insights we gathered from the event.

Interspeech.jpg

Using Large Language Models for ASR and SLU

A major topic at the conference was the use of Large Language Models (LLMs) to improve Automatic Speech Recognition (ASR) and Spoken Language Understanding (SLU) systems. Unlike traditional ASR models, LLMs are trained on large amounts of text data, which helps them better understand the context of spoken words. This makes them particularly useful for better recognition on domains with missing or scarce acoustic training data.

The research team (Jinlong Xue, Yayue Deng, Yicheng Han, Yingming Gao, Ya Li) showed how they are using LLMs to fix errors in ASR outputs, while others are exploring new methods for real-time speech recognition. There is also a growing trend of adding more text data to the training process for acoustic models, which improves their ability to recognize less common words and phrases. This approach is gaining popularity as it leads to better overall performance of ASR systems.

To learn more about how to create your own large language model, check out our step-by-step guide.

Improving the Whisper Model

The Whisper model and its updated versions were one of key topics at the conference. The research is focused on improving Whisper’s performance, such as making the text alignment more accurate, reducing errors, and better detecting pauses in speech. One example is Crisper Whisper, – a modified version originally designed for medical diagnostics. It offers more precise timing and fewer mistakes, making it a strong alternative to WhisperX for various uses, like legal transcription where high accuracy is needed.

The conference also highlighted several Whisper-like models using open datasets, which could make these advanced tools available for a wider range of applications. This growing ecosystem of Whisper-based models shows promise for improving accessibility and adapting to different industries and languages.

New Ways to Detect Fake Audio

As synthetic voice technology improves, it’s becoming more important to detect fake audio, known as deep fakes. The paper highlighted several new methods for spotting these manipulations. Some researchers are using advanced models to identify small differences between real and fake voices, even when the fake is very convincing.

These developments are crucial for keeping voice data secure in areas like security, media, and entertainment, where it’s important to trust that audio is genuine.

Better Speech Recognition for People with Speech Disorders

In the healthcare field, a major update was the release of a new dataset for dysarthric speech by Mark Hasegawa-Johnson, the creator of the UASpeech corpus. This dataset is designed to help improve speech recognition for people with severe speech impairments and will be used in a research challenge this November to test new models and approaches.

Google continues the development of its project, previously known as Euphonia, which aims to enhance speech recognition for individuals with speech disorders. They presented new techniques that make their models better at understanding and transcribing slurred or irregular speech, benefiting users with conditions like cerebral palsy or ALS.

These advancements are essential for developing technology that can accurately recognize and respond to the unique speech patterns of people with dysarthria — an area where personalized adaptation has shown dramatic accuracy gains.

Speech Technology for E-Learning

ASR (Automatic Speech Recognition) and speech technology in e-learning were well-covered topics at the conference. Many of the techniques presented were similar to the ones we developed five years ago. For example, using phonological features for mispronunciation detection. The paper is describing using a wav2vec2 model with modified CTC loss, while we used a transformer model for similar tasks as well.

What’s Next for Speech Technology

The conference offered valuable insights into speech technology's latest trends and future possibilities. We’re looking forward to using these ideas in our current projects and working with the research community to explore new possibilities.

We plan to use LLMs to improve ASR accuracy, adopt Whisper model enhancements for specialized transcriptions, and refine our speech recognition capabilities for individuals with disorders. For e-learning, we're exploring real-time feedback solutions for better pronunciation training. Beyond research, these ASR advances are already powering real-world deployments like voice-driven ordering in drive-thru chains.

Stay tuned for more in-depth reports on the topics discussed at the event. If you have any questions or would like to talk about any of these findings, feel free to reach out!

RELATED BLOG ARTICLES

View all Articles
How to Build Reliable AgTech AI When Farm Data Is IncompleteHow to Build Reliable AgTech AI When Farm Data Is Incomplete

In agriculture, missing data is part of the job. Farm data comes from different sources, under changing field conditions, and at different points in the growing cycle, so a complete and perfectly synchronized dataset is rare. AgTech models still have to work with whatever information is available. Some gaps barely affect the result, while others remove an important part of the signal. Knowing the difference is what makes the model useful outside a clean development dataset. Agricultural data is

# Agriculture
# AI / ML
# Data Science
Building Domain-Specific LLM SystemsBuilding Domain-Specific LLM Systems: When to Use RAG, Fine-Tuning, or Neither

When an Air Canada customer asked the airline’s website chatbot about bereavement fares, it told him he could book first and claim the discount within 90 days. The same answer linked to a policy page saying retroactive requests weren’t allowed. The passenger followed the chatbot’s instructions and later had his refund request rejected. The civil tribunal found the airline liable for negligent misrepresentation after concluding that he’d reasonably relied on the inaccurate guidance. The correct i

# AI / ML
# Data Science
# LLM
Scaling AI InfrastructureScaling AI Infrastructure: Navigating GPU Orchestration and Cloud Costs

Microsoft spent $37.5 billion on infrastructure in the quarter ending December 2025, with roughly two-thirds going mainly to GPUs and CPUs. At this scale, even modest waste is expensive. NVIDIA found that idle workloads consumed about 5.5% of GPU capacity across its research clusters. By combining GPU telemetry with job data and automatically clearing stalled or abandoned workloads, it reduced that waste to about 1%, potentially saving millions. Scale down the arithmetic and the pattern holds: a

# AI / ML
# DevOps
Forward Deployed EngineerWhat Is a Forward Deployed Engineer and How to Become One

Forward deployed engineer was once a niche title used mainly by companies such as Palantir. It now appears across OpenAI, Anthropic, Google Cloud, Scale AI, and other enterprise AI providers. FDE brings software engineering, technical consulting, and project delivery into one role. Its growth also shows what enterprise AI companies now need from engineers: an understanding of customer workflows, the ability to make sound technical decisions, and responsibility for moving a system into production

# Tech
# AI / ML