Speech-to-Text on Google Cloud is a tool used to convert speech into text using an API powered by Google’s AI technologies. The vendor states users can transcribe content in real time or from stored files; deliver a better user experience in products through voice commands; and, gain insights from customer interactions to improve service.
$0.02
per min
Phrase
Score 3.4 out of 10
Small Businesses (1-50 employees)
Phrase is a Language Intelligence provider. Its enterprise platform automates, manages, and delivers multilingual content. Global brands use Phrase across hundreds of languages to reduce time to market and deliver consistent brand experiences worldwide. The Phrase Platform brings together translation management, software localization, multimedia localization, machine translation, workflow automation, and language AI in a single integrated environment. From marketing…
$27
per month (billed annually)
Pricing
Google Cloud Speech-to-Text
Phrase
Editions & Modules
Speech-to-Text V2 API
$0.016
per min
Speech-to-Text V1 API
$0.024
per min
Freelancer & LSP plans - Freelancer
$27
per month (billed annually)
Developer Plan - Software UI/UX
$525
per month (billed annually)
Freelancer & LSP plans - Professional
$525
per month (billed annually)
Business Plan - Team
$1,245
per month (billed annually)
Business & Enterprise Plans
Custom
Offerings
Pricing Offerings
Google Cloud Speech-to-Text
Phrase
Free Trial
Yes
No
Free/Freemium Version
Yes
No
Premium Consulting/Integration Services
No
No
Entry-level Setup Fee
No setup fee
Optional
Additional Details
Speech-to-Text V1 API
V1 offers data residency for multi region only. Models include short, long, phone call, and video. V1 does not include audit logging. New customers get $300 in free credits and 60 minutes for transcribing and analyzing audio free per month, not charged against your credits.
Speech-to-Text V2 API
V2 offers data residency for multi and single region. Models include short, long, telephony, video, and Chirp. V2 does include audit logging and support for customer managed encryption keys.
Google low latency streaming api seems to be working best when compared to other cloud-supporting tools as this will help in realtime transcription for customer interactions. while comparing azure and amazon the google support more than 125 languages and ascents as we work with …
Earlier we were completely reliant on text pad or notepad, where we used to manually capture the information, which seems to be very hectic for the long-running meeting, because holding the information and capturing them and redocumenting it is very big process and it involves …
They just remind me of each other. Whenever I have a question, whether for personal or for professional reasons, I take out my smartphone, click the Gemini app, and then click the mic to ask my question and have the answer read back to me. I love Googles AI System.
Firefly is a great notetaker application that plus into meetings and organizes the data. However it comes presctured while the Google Cloud Speech-to-Text application you can better organized the data. Then have the option to plug it into other platforms to further organize the …
It delivered high accuracy in accented and noisy environments. Regarding its language support, it offers a variety of languages and dialects. Its's Api's are well-documented and easily integrated with our GCP-based stack. Also, its deployment is fast, and it is cost-effective. …
One major setback is the integration of multiple languages, where Google has support for more than 120 languages, and Amazon only supports approximately 30 languages. Regarding the transcription behaviour, Google Translate is very accurate, but Amazon Translate sometimes spells …
Descript is definetly less accurate than Google Cloud tool, while Google Cloud Speech-to-Text does have some troubles with overlapping and background noises, it still performs better than Descript. However Descript has some video editing functions which are not available in …
Otter is good for simple note taking and its UI is quite simple. But it seriously lacks in providing appropriate transcription as it is less accurate with Indian accents and being unresponsive in real time transcriptions. On the other hand Google Cloud Speech-to-Text excels in …
We use Google Speech to Text on the recommendation of a partner who uses it, in fact we do not evaluate other applications such as Amazon Transcript or similar
Google Cloud Speech to Text has a significantly cleaner and easier-to-use User Interface. If the user is already familiar with the Google Cloud product suite, then onboarding with this software will be an extremely smooth process. If a user has previously used other …
I like Google Cloud Speech-to-Text the most when it comes to other apps I have used so far. It have reduced my work, saved lot of time and made me less stress in meetings. It has also helped us in taking requirement gathering, knowledge transfer important notes to further …
While both Speechify and Google Speech-to-text do the job, certain elements that I find missing on Speechify are: it only works on Desktop with Windows OS, the customizations aspect is missing, there is no mobile app support (people these days want everything on their mobile …
I've also trialed IBM Watson Speech to Text for similar use cases. While both are highly capable, I find the Google Cloud Speech-to-Text software's accuracy and integrations to be a cut above. Harnessing Google's speech recognition prowess has elevated our firm's value …
Google Cloud Speech-to-Text outperformed its competitors significantly in terms of accuracy, surpassing any other product available. Additionally, its support for multiple languages was unrivaled in the market. Moreover, for clients with robust bandwidth, Google Cloud …
1. It's an efficient tool for improving efficiency by saving a lot of time in typing. 2. It saves at least 40-50% of our time, thus increasing efficiency. The amazing thing I liked about it is the accuracy with multiple accents & multiple languages. 3. It also takes …
The accuracy of Google Cloud Speech-to-Text is much better than any other tool. It has better API integration with 3rd party tools. The transcription is on at real-time basis with the best efficiency. It has good language support from across the globe. It provides better noise …
Google Cloud Speech-to-Text is better than these other services. The main driver is the cost for the service and what you get, the value proposition is very good. Also, the scalability of Google Cloud Speech-to-Text is great, so that down the line, as our needs change and …
It is very expensive and more difficult to navigate. Compared to the other support is slow and ineffective. Dealing with anything more involved is all but impossible. It has become more prone to bugs since being bought by TMS. Support suggests things like "use incognito" mode …
I did not select Memsource. I use it because my clients use it, but I find it very useful especially due to a friendly interface and quick jumping between segments of different statuses during translations.
SDL Trados Studio is more robust but also much more bloated and resource-intensive and ultimately less flexible. I use both daily, but Memsource is my go-to in almost every situation.
Memsource is quicker and easier to use. You can even start translating on your mobile phone or tablet, as long as you have an internet connection. Plus Memsource shows me a live preview of my translation. I think Memsource is great for beginners as well, since you don't need a …
Memsource is the best, because it does not have any unneeded functionality that would compromise its intuitive use, and those functionalities it includes are all easily accessible.
Memsource stands out because it can be used in ANY browser. Thanks to its mobile app, it is easy to accept and review the work that translators are assigned.
I also use Wordbee, but the thing I don't like it most about it is that the number of segments you can display at a time is limited to 100 and you must move around many pages if the job size is large. I use SDL Trados Studio as well, but not cloud-based, so not really …
Memsource is the easiest to use among all other CAT tools out there. Even when you use it for the first time, it will be easy for anyone to use. You don't even need much instruction or a tutorial process. Still, it allows you to do a lot. It's easy to find some features you …
There is no true comparison because Memsource consistently outperforms the competition because it is simpler to use and has greater TM and MT integration.
Memsource is way cheaper and easier to use. Other softwares seem really old-fashioned and out of date. They are very difficult to use for a new learner and the UI is not good looking. Also, the price for products like Trados is too much expensive for a new freelancer and many …
Memsoucre has a more friendly user interface. It keeps offering new features (like machine translation add on) which wasn't there when we first started using Memsource. Very smooth and easy communication with customer service team. My tickets are addressed quickly and …
I have been using memoQ and SDL Trados Studio for a longer time, but the Web editor for Memsource is the easiest to use of all. I do like the functionality in memoQ better, but Memsource is easier for our translators.
What distinguishes Memsource from other translation tools is that it does not require massive specifications to run. Average computers can run it smoothly without needing 8 or 12 Gb ram. Being able to use the browser to do the translation task is also another excellent feature. …
Memsource is the platform chosen by one of my clients, Booking.com. I have direct contact with my client, and the machine translation service offered is better than the one offered by its competitors. It's more intuitive and easy to use as well.
Memsource includes all the basic features a freelance translator needs in a very user-friendly approach. All you need is a click away and provides quick results.
Memsource has a very good user interface that people can quickly learn and start using. Also, its analysis features are top class which helps me provide a detailed estimate to my client based on the file particulars. It has a QA feature that helps me do a very high-quality …
Memsource is more user friendly and innovative. The interface looks better and more accessible. I have not experienced any downtime with Memsource. The times that I used Memsource it didn't give me any errors or issues. But with the other tool, there were several times when …
So, I've had scenarios like when I collaborate with a team where the people are from around the world. So, I used it there, and we spoke to each other in their native language. That boosts everyone's confidence in our collaborative efforts. I've also utilized its model and the API in my projects, including a Virtual assistant and a multilingual application that allows us to learn languages from around the world. We tested it with a group of 12 people, and that's when it failed. I mean, it's not a failure, but it can't detect every person.
I like Memsource because when I am not home and I don't have my laptop, I can borrow a computer, log in to my Memsource account and I'm ready to begin translating. I can even download the source files and check the TM and glossary there. It's not necessary to download the software and lose time on that.
In the case of EN to JA translation, we enter text and then convert it to get the correct final text (word or phrase) which uses correct Chinese character(s), as there are often multiple Chinese characters with the same reading but different meanings. When those conversion options are displayed, it is usually possible to change our selection among them by hitting the Tab key, but in Memsource, hitting the Tab key makes us leave the text conversion and move to the source segment, so we must make sure to use arrow keys to choose the right conversion option when using Memsource.
When there are tag elements used in the source text, the tags must exist in the target text of course, but the tag order also must be the same. The tags cannot be moved around in the segment, which causes problems in the case of the Japanese language, because the word order differs between EN and JA.
In the Japanese language, Italics are basically not used, so there must be no texts between the tags which specify Italic font. But then the segment cannot be confirmed and the job status cannot be changed to "Complete".
I use and manage an Academic edition (approx. 15 students) and the corporate license (10 users). In both cases, the features are easy to find, set and use. My students and my translators understand the dynamics of the platform easily and get used to them quickly.
The reasoning behind my 10 is that the UI is very intuitive; I didn't require any formal training to use it. Google's speech-to-text is not just a conversion tool; it helps automate mundane tasks, saves time, and has an almost human-like understanding.
It is designed to be a lean system, but that means that things are clustered in weird groups and a lot of trial and error is involved getting the settings right. If something goes wrong first level support is very quick and has very fast answers which are quite typical of first level support and essentially equate to asking you if you have switched it off and on yet. Anything more involved is challenging for support to address.
We rarely use support, but most questions were answered in a timely fashion, although we didn't exactly find them satisfactory. That's mostly the fault of the software and not the Support team because we asked for things that Memsource couldn't do.
Earlier we were completely reliant on text pad or notepad, where we used to manually capture the information, which seems to be very hectic for the long-running meeting, because holding the information and capturing them and redocumenting it is very big process and it involves human work we were looking for some automation which can fix this issue then we got Google Cloud Speech-to-Text which converts audio to text files easier and faster and also it support many languages there by helping to align with various different clients across the globe and make the discussion seamless.
It is very expensive and more difficult to navigate. Compared to the other support is slow and ineffective. Dealing with anything more involved is all but impossible. It has become more prone to bugs since being bought by TMS. Support suggests things like "use incognito" mode to log in. Use a different browser to log-in. That is a concern given the customer data involved if things at that level are not being fixed. If that's what it's like in the showroom, what is the kitchen like?