AI in Azimuth Software Solutions
At recent industry events, we showcased how our clients utilize AI capabilities integrated into Azimuth Soft TV production automation software.
A major customer uses our TT software suite, to prepare subtitles for broadcast content. The latest version of the product’s subtitle editor introduces AI-powered speech-to-text functionality. Editorial staff import video or audio files into Azimuth Subtitle Editor 3.2 and receive an AI-generated transcript of the audio, which serves as the basis for quickly and efficiently creating broadcast-ready subtitles.
Although current technology means recognition results aren’t always perfect, our users report significant gains in speed and workflow efficiency when handling large volumes of material.
Our AI engine is constantly evolving; clients already have access to features such as automatic punctuation and the conversion of text to numerals. While these tasks may seem simple, they require a certain level of AI maturity to be used effectively within a high-paced production environment.
It is crucial for our clients that their content remains within the broadcaster’s secure production network. Consequently, our solutions rely on a local AI engine and do not require Internet access to operate.
The potential for using AI to enrich metadata for media assets in TV production is virtually limitless. In the near future, we can easily envision AI that automatically detects scene changes in video files, flags and blurs on-screen actions unsuitable for broadcast, and provides detailed descriptions of objects and events to facilitate MAM search.
Nevertheless, we offer our clients only those AI features that are in genuine demand and can deliver tangible benefits to the broadcaster right away. The next step in this direction is a new version of our file converter, FileSynchronizer. Now, when importing files from various sources into our APX MAM system via a watchfolder, the new clip will include—alongside the production-quality and low-resolution video files—a metadata field containing an AI-generated transcript of its audio. Newsroom staff, who constantly handle large volumes of video files in a wide variety of formats, will particularly appreciate this feature.
If the source format matches the standards used at the facility, the video is imported without re-compression at speeds of up to 10x real-time. The speech recognition process can be further accelerated by outfitting the conversion servers with graphics cards.
Having a video transcript with timecode markers for spoken phrases within the MAM system will also allow us to implement a “text-based editing” workflow in our NewsBase video editing module—a feature highly popular among users of the latest versions of Adobe Premiere.


