There was a time when I was working on a text-to-speech system for our automated trading program. The voice we got back was either too slow, making listeners drowsy, or too fast, sounding rushed and unnatural. Our team tried many approaches, but every time we implemented it in a real-world scenario, we encountered problems. Today, I want to share how we reached the point where our Thai text-to-speech sounds natural, and what we learned from our mistakes.

The Problem with Unnatural Speech

Initially, I and my team tried using existing systems to generate speech for gold and forex trading data. However, the resulting voice was truly unsuitable. We tried various speeds, but we always encountered issues. When the voice was too slow, listeners would become drowsy because trading data needs to be read continuously. But if the voice was too fast, it sounded unnatural and wasn't suitable for reading financial data that requires clarity. Sometimes we received feedback from customers saying the voice sounded like a robot reading, which was something we needed to fix quickly because it affected the user experience of our iCafeFX app.

Initial Testing and Measurement

To solve the problem of unnatural speech, we started by measuring the reading rate in characters per second (CPS). We broke down text into segments and measured the time it took to read them. The first segments we tried were 338 characters and another of 403 characters. We adjusted various values and measured the results, using a system we developed on a Mac Mini M4 Pro with a private NAS for data storage and an RTX 5060 Ti for processing. The early test results showed we hadn't yet found suitable values for real-world use.

Initial speech experiment
Initial speech experiment showing unnatural results

Finding the Optimal Value

After the initial tests didn't yield the desired results, we began searching for an optimal CPS value through detailed experimentation. Based on the actual data we measured, a natural-sounding rate was approximately 13.6 characters per second, which was much lower than the values we initially tested. However, we needed to ensure this value would work long-term. Therefore, we conducted additional tests using larger texts and made detailed adjustments. We found that the optimal value for our system was in the range of 14.51-16.51 characters per second, a rate that can be practically used in daily operations.

System Improvement

After finding the optimal CPS value, we began improving our system further. We used the FLUX model to enhance voice quality and Redhat MCP AI to improve processing speed. We needed to address the potential issue of reduced voice quality when using lower CPS values. Additionally, we had to fix inconsistencies in certain parts of the speech, which we resolved by adding supplementary audio processing. This made our speech sound more natural and consistent, especially when reading complex trading data.

System after improvement
Improved system showing natural speech results

Real-World Implementation

Once we had the optimal CPS value, we implemented it in our MetaTrader 5 system, which we use for gold and forex trading. Reading trading data with natural-sounding speech helps traders follow information more easily and reduces fatigue from manual reading. We received feedback from users that the speech sounded much more natural, making our system more comfortable to work with and helping to reduce errors in understanding trading data.

Lessons and Next Steps

From this experience, we learned that small adjustments can significantly affect user experience, especially regarding voices used in systems that run continuously. We also learned that detailed experimentation and measuring results based on actual data is crucial. In the future, we might develop the system to automatically adjust CPS values based on data type or user preferences to provide the best possible experience.

For those who want to try real gold trading, open an XM account at: open an XM account via our partner