The Future of Rhasspy

Hi everyone,

As some of you may know, I recently changed jobs and joined Mycroft AI as a senior developer. The two reasons for this change were (1) I needed a fully remote position in order to move closer to family, and (2) my previous job was going to force me to stop working on open source. No thanks!

Since I’m not independently wealthy :face_with_monocle: , I applied to a few places and Mycroft made me a good offer :slight_smile:
Obviously, this begs the question of what is going to happen with Rhasspy, so I’m here to talk about that.

Most importantly, I am not abandoning Rhasspy. However, updates are going to come slower due to time constraints. Mycroft hired me to help get their Mark II smart speaker over the finish line, and I am going to do that. I’ve given Github access for Rhasspy to a few folks already, but if any other maintainers would like to volunteer, please do :raised_back_of_hand:

My long term goal is to get many of Rhasspy’s key features into Mycroft. As you can see, I’ve already got offline speech to text and text to speech working! Satellites and Hermes protocol support are something I’d also like to do down the road.

Ultimately, I hope our communities can join together at some point. I stayed away from Mycroft initially because of its dependence on a cloud server, but that’s going to change now with me there. Other things, like a skill store and a dedicated hardware device, I just didn’t have the time and resources to produce. So it would seem both sides here would have something to gain.

What are your thoughts?

17 Likes

Congratulations with the new job!

I stayed away from Mycroft for the same reason as you did, and I’m happy to hear that its dependence on a cloud server is going to change with you as a developer there!

What are some things we in the Rhasspy community can do to ease the path of joining both communities together?

2 Likes

Congrats on your new job!

I have taken a look into Mycroft a while ago but the main reason for now using it was the online recognition. Also, they did not support Dutch.
You joining them gives new hope to that :smiley:

Hopefully I can put some more time into solving issues for Rhasspy, but my spare time is also limited.

Keep us up to date and good luck!

2 Likes

Nice to have some news and congrats for your new job, you highly deserve it !

I never dig up into mycroft due to cloud so I don’t how it could and even would replace rhasspy. Custom wakeword, home automation intégration, custom intents and such comes to mind …

Really hope rhasspy will continue to be enhanced for a while.

3 Likes

Hi !
I would say, as far as you keep your open mindset and still beleive in Rhasspy roadmap, it’s a triple good news:

  • You find a job to fullfill your essentiel need (pay the bills !) :money_mouth_face: in relation to your passion
  • You may bring Mycroft out of the dark clouded side and keep them far away from the GAFAM :angry:
  • Rhasspy audience may rise again and then get new supporter and may be contributors :nerd_face:

Cheers and all the best in your new job (happy Mycroft people)

2 Likes

Thanks for the kind words, everyone! I’ll keep you up to date on my plans; please feel free to message me with ideas.

For me, it would be helpful to get a sense of what people find most useful about Rhasspy, and where Mycroft is lacking. I already gave my own examples, but the Rhasspy community is pretty diverse. I don’t want to leave anyone behind, but we will need to make some decisions about what to “take with us”.

Congrats on the new job!

Coming from someone who’s done a lot of integration with Rhasspy for Home Intent, I can definitely say your technical documentation has been amazing. All the ins-and-outs of hermes and setting up various configurations (satellites!). I’ve also really enjoyed that all the various components are pluggable, easily controlled via the API, and it’s all listed as to what works with various languages. I know you’ll bring a lot to the table at Mycroft and they are very lucky to have you!

I’ve also been very impressed with your work on Larynx, and hope that it will continue!

1 Like

For me, language support and offline STT is key.

3 Likes

having worked a lot on the use of the Snips application before its takeover in early 2019 by the company Sonos, and during the shopping of Christmas gifts I was able to observe that the sonos speakers currently only worked with Google home and Alexa I conclude that to eliminate the competitor Snips Amazon and Google had to deal with sonos to definitively eliminate the Snips application. Which app is closest to the offline aspect of Snips right now ??? What do you think of the Alice project compared to Rhasppy? I have worked on other projects during this time and I have to get back to them. Thank you for your advice Regards

1 Like

Rhasspy

You can find a topic about that here:

Congrats on your new job.
Long live to both Rhasspy and MyCroft.
I think both can run offline! People have used MyCroft offline in packages for Kids and places where internet access had been issues even though not official.

I will be happy to see both Rhasspy and MyCroft grow more and more… even collaborate together

1 Like

Congrats and I understand why you did this. But I see this as a big loss, similar to when Snips was acquired. I think it means the death of the Rhaspy project even if more slowly than Snips. *** I hope I’m wrong. *** The problem is when companies take over, almost always, the open source aspect goes down. Meaning the ability for programmers to use and add onto a tech platform without friction - obtaining/paying for licenses, etc. I can see this as a win for the open source, voice developers, community if two things happen… 1) Mycroft eventually gets all the Rhaspy functionality you will hopefully help give it. and 2) Mycroft embraces the OS/developers and doesn’t go the same path as Sonos.

1 Like

I’m happy to hear that you have a positive attitude toward both Mycroft and Rhasspy.

There is something of a history of voice computing projects losing steam after developers are hired by commercial projects. The Jasper project died after the developers were hired by Microsoft to work on Cortana. KSimon has been limping along since the primary developer was hired by Apple to work on Siri. So it’s not without reason that people were concerned.

I have also avoided Mycroft for years because I don’t buy their marketing around the cloud component (“we are open source, so you can totally trust us”, “using our servers helps anonymize your activities”). I’d love to feel more comfortable inviting that project into my personal space.

I’m happy to see that Mycroft does make the Selene back end available. I will try and set that up again (the last time I tried was several years ago using an unofficial and unsupported repository, and I gave up after a couple of days). One place where I do see the benefit of a centralized back-end is for organizations like schools and health care providers who need to keep PII protected while also providing a consistent and custom set of services.

Congratulations and best of luck with your new position. Thank you for the efforts you have devoted to open source voice interfaces.

1 Like

Thank you again to everyone for the warm replies!

Very happy to hear this :slight_smile: As @romkabouter mentioned, language support is very important both with (offline) STT and TTS. Every language we can get enough data for is another group of people that don’t have to rely on Google, etc.

It will! I’ve already started on a follow-on that I’m calling Mimic 3 (to fit Mycroft’s naming scheme). It’s close in architecture to Larynx, but is using a new TTS model that is almost 2x faster on the Pi 4. I’m still figuring out the best hyper-parmeters to use, but once I do I will retrain all of the voices and upgrade Larynx :+1:

I understand, and I won’t pretend like I can 100% promise what the future will hold. However, at least with Mycroft what I work on will be completely open source. Snips’ tech was amazing, but it’s effectively worthless now because of the Sonos acquisition.

I will be working to get Rhasspy’s tech into Mycroft over the coming year. The Mark II will hopefully be shipping in September (if they get enough orders), and my goal is to have offline STT and TTS for all of Rhasspy’s supported languages on board.

Definitely, and unfortunately. I feel like having some shared standards could help ease the burden of shifting between projects. For all its warts, Hermes enabled a lot of Snips users to switch over to Rhasspy without redesigning everything. I hope to finally have some time to work with @sepia-assistant and @DANBER to establish those standards.

No company is immune from acquisition, so you shouldn’t trust them! Mycroft seems very unlikely, however, to go down the same road as Snips.

You’re welcome, and I see myself as a long ways from finished :smiley:

1 Like

Congrats on the new job!!!

Seeing everything you contribute will be open source, what works for one project will most likely work for the other project, so win win.

1 Like

For me Mycroft is lacking three important things :

  1. TTS Support is very limited. For German I only have the unusable slow Mozilla Deepspeach Model or low quality options. Hearing you want to create an even better Larynx called Mimic 3 which can be used with Mycroft sounds really great.
  2. Im not using Home Assistant for controlling my Smart Home, so I need an option to process the intend in NodeRed or at least configure Get & Post HTTP API Requests. And it’s just great if you want your own automation flows. Like my alarm flow im currently building in rhasspy does not only ring an alarm at some time but also powers up my room heating and slowly increases brightness of my lights beforehand. (While writing this comment i found this : https://github.com/JarbasHiveMind/HiveMind-NodeRed . So there seems to be some kind of community skill which is able to do this)
  3. It couldn’t be used completely offline. Sounds like this is going to change too.

Im actually really looking forward to this.

5 Likes

One trick I use to improve performance is to use the engines with caching.
So what I do is have most responses as simple as possible, “like the light is on” rather than “the bedroom light is on” and for things that are more dynamic like weather, temperature, humidity, time date, etc I actually have node read send the sentence to TTS whenever it changes but only play it when there is a request, so it is already cached and the response is close to instant.

Jarbas is an excellent buddy with one stop shop for everything AI, Voices and more… :smile:

Happy Adventures with the new job.

For me Rhasspys’ ability to do custom intents is my favorite tool.
My Rhasspy installs: In a silly bot, in two solar monitoring stations (RV and house), in an offline house assistant, in a offline RV assistant, all simply give info to my custom python programs which do all the handling, and send responses back, go online and post this or that api, do gpio stuff, turn camera on and off whatever.
Rhasspy has the Power of Whatever: so that’s my favorite feature.
I did play with its home assistant integration as well, but found my custom python stuff much more usable for my unique projects.
For me to move to mycroft:
*. No need to be online, unless updating something, and no need to have anything posted to someone elses’ cloud server to do anything, sure If I need the weather I can bring LAN or Wifi up (I have an intents for that), Python does its thing like a request to a weather api, then LAN down wifi down (intents for those as well). It all works with Rhasspy.
*.I do use the mycroft custom wake word (I log into an older install of Ubuntu Studio I didn’t erase just because I setup the mycroft custom wake word tools on it awhile back.I’ve not needed to make any new ones in awhile though.

2 Likes

I deeply respect you @synesthesiam. While It saddens me that you’ll surely have less and less time for Rhasspy, I understand the need to eat and pay the bills. What devestates me far more than losing Rhasspy though, is that your skills are being gained by somewhere as deceptive as Mycroft.

Yes they tout loudly that their code is open, but this is not a privacy respecting project, nor is the company behind it. I came to Rhasspy precisely to get away from the likes of Mycroft. There are reasons that their project is specifically structured to default to using their cloud services (sound familiar?). There are reasons they demand user logins (also familiar…?). Their project has been specifically structured in a way that makes self hosting, and I quote “not easy and is unlikely to provide an equivalent user experience [as sending your personal info to their computers].”

I mean, is Mycroft different than GAFAM? Maybe, but if so, it doesn’t seem to be for lack of trying to emulate them. Let’s take a quick look, not at their glossy marketing homepage, but at their privacy policy (which covers both their website and their services). Just a quick skim through, copy and pasting the privacy relevant bits comes up with:

When you use our Services including the Mycroft Voice Assistant, your voice and audio commands are transmitted to our Servers for processing.

Because they made it, you know “not easy” to self host, like say Rhasspy currently is.

We collect information about you directly from you and from third parties, as well as automatically through your use of our Site or Services

Just to clarify that it’s not just “opt in” stuff.

When you use our Services, your audio commands are transmitted to Mycroft for processing, as part of the Services. We may also collect other metadata about your audio commands, such as the time and location.

Note that it’s not limited to information needed to fulfill the request.

we collect information about your device, including platform type and location

Once again, specifically not limited to fulfilling the service.

If you comment or post content to the Services, we may gather data about the content you post.

Just in case you were wondering if they would also use the content.

We automatically collect information about you through your use of our Site and Services, log files, IP address, app identifier, advertising ID, location info, browser type, device type, domain name, the website that led you to our Services, the website to which you go after leaving our Services, the dates and times you access our Services, and the links you click and your other activities within the Services.

That’s a list that would make Google proud, especially the advertising ID. What could that be for?

To send you news and newsletters, special offers, and promotions, or to otherwise contact you about products or information we think may interest you.

Because privacy focused companies are all about promotions and special offers.

for other research and analytical purposes

(how much vaguer and more all encompassing can you get than “other research” and “analytical purposes”?!)

To protect our own rights and interests, such as to resolve any disputes

So you control the assistant in my home and can use it’s info against me in a legal dispute?

We may share your information, including personal information

Just to be clear…

We may disclose the information we collect from you to third party vendors, service providers, contractors or agents who perform functions on our behalf.

Once again, pretty much to anyone.

If we are acquired by or merged with another company, if any part of our assets are transferred to another company, or as part of a bankruptcy proceeding, we may transfer the information we have collected from you to the other company.

As with pretty much any company… which is one of the problems with company run projects.

We may share aggregate or de-identified information about users with third parties for marketing, advertising, research or similar purposes.

Just in case you thought they were only collecting to improve the project.

We and our third party service providers use cookies and other tracking mechanisms to track information about your use of our Site or Services. We may combine this information with other personal information we collect from you (and our third party service providers may do so on our behalf).

Oh? Like who?

We use automated devices and applications, such as Google Analytics

Ah, that’s who (at least one of the who’s…).

our Site does not recognize browser “do-not-track” requests

Real “privacy focused”, right?

If you’d like to update your profile information with us, you may do so through your account. […] we may maintain a copy as part of our business records.

So you can ask us to delete it, buuuut… “business records” y’know?

Our Services are not designed for children under 13; and children under 13 are not permitted to have an account with us. If we discover that a child under 13 has provided us with personal information, we will delete such information from our systems.

Why no users under 13? Much like Audacity moving to ban users under 13 when they added tracking and Google Analytics. In many areas it’s illegal to track personal data about children for commercial purposes. If you want an easy surveillance advertising based income, you gotta get rid of users under 13.

A couple of these by themselves could be understood, but what’s the overall pattern here? Is this the pattern of a privacy centric company, recording advertising ID’s, location, IP addresses, content and using it to give to third parties and marketing?

I’m sorry if this doesn’t come across as supportive as the other comments here. Most likely this post will lead to personal attacks against me by people who feel that they’re defending you, all while encouraging you towards being controlled by “© Mycroft AI, Inc.”’. Yes, the code may still be open, but make no mistake, this is open code to use against privacy. If I didn’t care about you I’d just be silent. I’m posting this precisely because to me this looks like a good, ethical, skilled person being bought by a company that has long fed off misleading FOSS and privacy advocates. This is a clear attempt at swallowing the more private competition, ie you and Rhasspy.

I wish you well, I really do, but this is a very sad day, indeed.

Hopefully a more honest company makes you a better offer soon. Hopefully I’m wildly wrong. Reading their privacy policy though, I doubt it, so I do hope you’ll be careful to keep an escape plan handy in case you need it later. I would have happily thrown hundreds of dollars at a crowdfunding campaign for Rhasspy, but Mycroft will never get a single cent from me.

Anyway, thank you for all the good that you did for Rhasspy and the community. It was good while it lasted.

5 Likes

I hope you are wrong as well, I was not using Mycroft but reading your post makes me think twice (and some more)

I will be sticking to Rhasspy anyway and I hope it will still evolve, maybe slower or maybe other developer pitch in.

Congratulations for your job ! This should help you for a better life…

Thanks for all your works on Rhasspy, very usefull and really private by design.
Things I really appreciate : fully offline tts and stt, docker installation, custom wake word (I’m not fond on “ok Google/Alexa/jarvis/my lad…”), ability to change speech recognition (deepscpeech is the future, well, i Hope so).

Best regards,
Damien

2 Likes

Thank you for the reply, @VoxAbsurdis. I’ve only worked at Mycroft now for a month, but I may have a little more insight than I did previously.

A lot of the problems you mention all seem to stem from the same issue: Mycroft’s (current) dependence on Google for their speech to text. If they send audio to Google (aggregated at least, but still), then their privacy policy must be at least as worse as Google’s. My hope is to get them off Google, and they are all for that :slight_smile:

As far as users under 13, I actually totally get that. “Child” is legally defined differently around the globe, so some cutoff us surely needed (not sure who came up with 13). I learned about this problem when I found out how bad the open source face recognition software was for photos with my kids: there’s almost no training data. Turns out, anyone who collects a bunch of data on children (pictures, audio, etc.) is usually not doing it for the good of humanity :frowning:

I had considered this, and I appreciate the offer; do you have any idea how I would sustain such a thing? I don’t really have anything to offer as a subscription (unlike, say, the Home Assistant folks at Nabu Casa). What would I have asked people to crowdfund?

I came across the Libre Endowment Fund, and thought something like that might have been a good fit. If anyone knows about something similar – a public fund for open source software – or a way to fund the development of open voice software for people with disabilities, I’d love to hear about it.

There are potentially a few options for open source crowdfunding for something like Rhasspy. From the simple GitHub Sponsor or Patreon approach (where you could have a devblog for folks who contribute). I’ve also heard of some folks doing a Kickstarter for a “year’s worth of development”. You don’t necessarily need something additional to offer, as the folks who would contribute likely just want Rhasspy, as it’s established and you’re always keeping it up-to-date (not to mention writing software that also has research benefits). Some folks have also been successful with getting grants for their open source stuff, but you’d have to find them first.

The downside would be increasingly more overhead (GitHub the least and grant-writing, likely the most) and getting up to a full developer salary would be tough. It’d likely require a bit of self-promotion. I’m not sure if we’re ready to value open source software properly as a culture, but I hope we get there.

2 Likes

Yep, that was my read from the start. And sadly it’s been the norm in the speech recognition industry for decades now. Every ASR/TTS platform that’s any good gets bought by on of the GAFAM gang members so they retain control over the industry - and the effect is higher prices and more difficulty for independent developers to build voice apps - which is probably their main goal… to make sure only the big gorilla companies own the voice-apps. My guess is Mycroft’s goal/exit-strategy is to get acquired by a gorilla. And because of this, I don’t think any dev should waste time building apps on/with it, otherwise you’ll be wasting your time like the devs who built on Snips.

3 Likes

This reminds me of Jeff Geerling, who I follow on a few platforms. He seems to be doing great, but I don’t think I could stand doing all of the necessary social media promotion in addition to actual development.

Absolutely agree. One person can make such a difference, but for some reason there is very little support. In my time working for the U.S. government, I saw millions of dollars wasted (in my opinion) on projects that, even if successful, got us no closer to something actually useful.

Just to be clear: I went to Mycroft asking for a job; they didn’t come to me. I interviewed with 4 other companies as well: 3 of them said that if I were hired, Rhasspy was to be shut down immediately and everything I did going forward was going to be closed sourced. 1 company was currently open source, but the backend they planned to create was not going to be (remind you of Snips?)

I get the cynicism, but I’d suggest looking at the broader picture. Even if Mycroft gets acquired at some point in the future, everything is still open. Contrast this with Snips, where they published an amazing paper about their training backend, but it was only ever a promise that they would open source it – and they lied!

I still believe our two best defenses against the GAFAM gang are (1) don’t support any company that isn’t fully open source and, (2) build interoperable standards between voice assistants so that when they inevitably do get acquired or abandoned, we can just shift to a new project without starting from scratch.

8 Likes

@synesthesiam Congratulations on your new job. Topic wise it seems like the perfect fit.

Rhasspy is quite impressive considering that it mainly was made by one person. I have always been impressed buy the amount of work and care you have put into it. Thank you.

But it was also sad to see how little traction Rhasspy was getting. There is a huge OSS community for smart home (Home Assistant and others) and a lot of those users are using Google or Amazon for voice control so I hoped more would switch to Rhasspy. The end result is: It still very much depends on you. there are just noch enough users and therefor developers.

I hope that making Mycroft more OSS and 100% offline usable works out. Until then I hope Rhasspy lives on.

(I would have been willing to support it on Patreon btw. But I don’t think there are enough users to pay you a fair wage using Patreon)

What I would need to use Mycroft are: 100% offline usage, German TTS and Hermes or similar open usable protocol (I am using it to show feedback to voice commands on LED matrix screens).

My naive plan/hope/strategy would have been:

You’ve mentioned Nabu Casa: I think Rhasspy fits very well in their open smart home vision. They don’t really have a working offline voice assistant solution in HA. It would be great if they could support a few months of work on Rhasspy with a focus on out of the box user friendly usage with HA. (integration in the HASS OS, integration with entites and auto generation of slots and some form of repository for intents/automations/scripts). With that done they could market Rhasspy as default offline voice solution for HA. Rhasspy would get more users from the HA community and more users means more supporters (devs and possible Patreon supports). (but well the other issue is: Readily available hardware without much DIY work that you can put in a normal living room. (basically a Pi with a speaker and good microphone(s) in a nice looking case.) )

In any way: Thank you for your work on Rhasspy and good luck on your new job.

1 Like

It’s fun that you mention that, as that is what I’ve been working on with Home Intent. Intended to bridge the gap between HA and Rhasspy by auto-creating slots/sentences/responses where users can just connect it to Home Assistant and it manages Rhasspy for them.

In the next month or so, I’m planning to get better satellite support and start looking into Hass addon store.

1 Like

Congratulations Michael on the new job - it’s great you are working at something you love, getting appreciated for it financially, and closer with your family !

I’ve had a quick look at Mycroft, and am also of two minds. Both projects will benefit from a closer association. I just hope that it continues the way you hope.

What is your relationship with Nabu Casa ? If they can dedicate a programmer to ESPhome (15% of their user base), then why are not doing more to provide offline voice assistant to the 75% who currently use the big commercials ?

1 Like

Oh nice. Didn’t know about it. Will give it a try when If ind some time.

Just a suggestion: I think it would be gerade to add somewhere early on the site that Home Intent is based on Rhasspy (and a link to this community) and is for use with Home Assistant. Or that it is the bridge between Rhasspy and Home Assistant or something like that. As it is now I would have thought “oh a new alternative to Rhasspy” if I randomly found the website.

Do have any plans to integrate something like “Apps” oder Scripts for functions that are more then just Home Assistant Intents? (Maybe just add a simple ui to enter python scripts that are using some the APIs developed by others in this community).

Edit: I should have read more in the documentation before askig about the script/app issue. The Component feature is for that. Then my question would be: What about CustomComponents and some form of repository for those.

Ah, I do actually mention it right away on the GitHub page. I’m actually not sure where people would head to first, but yeah, I’ll update the homepage of the docs to indicate it’s based on Rhasspy!

I do like the idea of a repository of custom components that are accessible via the UI! I don’t know of anyone who has written a custom component yet, but it’s definitely a consideration for down the line.

My pet beef with Mycroft was an uneasy feeling it was the same as it presented itself far beyond of what it was capable.
Its only in the MKII that they actually process the incoming audio with DSP echo cancellation and beamforming and Rhasspy suffered the same as really in the presence of noise it was unusable.
The delay and cancelled crowdfunders and this thing of ‘patent trolls’ just had my alarm bells going.
They have a new CEO and we will have to see which way it goes but there is a strong possibility you could be right.

synesthesiam like all needs $ and has to work and if you have got to work you have to work.
I am not a fan at all of Mycroft because to me it still seems the goal is to present the semblance rather than the real thing and hence why I am dubious.
As for working for them his core skills are bang on the button and the guys got to work even though I am not a fan of the company as likely 99% of us are not doing what we want but doing our best to earn $

I have told synesthesiam what I think of Mycroft but don’t and shouldn’t have any opinion on what someone needs to do to earn $

2 Likes

I’ve spoken with Paulus (creator of Home Assistant) a few times. He definitely sees the value in local voice assistant tech for HA. I’m hoping to collaborate with Nabu Casa through Mycroft; who knows what will become of that :wink:

Looking at the history of Mycroft, I agree. They were talking about full-on conversational assistants years ago, and the software they have today is nowhere even close to that. The new CEO seems much more down to Earth, though he’s still forward thinking.

On this topic, I think people may not have realized that my last job was for the U.S. military. So even with the many reservations folks have about Mycroft, it’s way better than the alternative! I’m grateful to my previous employer for getting me off to a good start, of course, but my work would have ultimately been shut down or relicensed.

6 Likes

Hi @synesthesiam, first of all, congratulations on your new job. It’s always good to work on something you have fun with:)

Would be great to continue the concepts we started in the other thread.

If you can work with a non commercial only restriction, I might have a nice dataset source for your trainings.


Regarding your questions about missing features in Mycroft, for me the most important, which also were some of the reasons to build Jaco, were:

  • Missing full offline capability
  • Problems with recognition accuracy because of the missing language model adaption
  • Similar to Snips the skills did only support python and not arbitrary languages and dependencies

And now in comparison to Jaco it’s also missing the skill privacy concept and some of the modularity.


By the way, it would be great if you could add Mycroft with your new STT approach to the benchmarks: Jaco-Assistant / Benchmark-Jaco · GitLab

Greetings, Daniel

1 Like

I thought your announcement was a bad dream when you first made it, but alas it is not.

Seems I’m a bit late to the game again, first it was Snips before the Sonos purchase and now Rhasspy. I guess if you end up shutting Rhasspy down at least what I have now will still work.

There are a few things that I don’t like about Mycroft, first is the obvious it’s cloud-based. From what I’ve been reading it appears that when compared to Siri, Mycroft does less to protect their users personal data than Apple.
The second part is the terrible detection Mycroft has for wake words. Now to be fair the last time I had Mycroft running was nearly 2 years ago but the experience was terrible. Snips detected everyone in the house for ‘Hey Snips’. There wasn’t a blasted thing I could do to get Mycroft to detect my wife or kids when they wanted Mycroft’s attention. The FAF (Family Approval Factor) was zero which lead me to Rhasspy.

I hope you will have Sway with Mycroft but I fear it won’t be enough and we’ll loose possible the best offline assistant around.

1 Like

Can we sadly say that we will never see Rhasspy 2.5.12 ?

After Snips, and maybe now Rhasspy, which is still far better than anything else, I ask myself going google/alexa route with all sadness to not revive one more time having to redo everything … :cry:

Of course Rhasspy still works nice, but if you are standing still, you are actually going backwards …

Congratulation for your new job !

It’s seems very positive, mycroft will probably reuse some part of Rhasspy, and empower it.

Mycroft has a business models plus people paid for working on voice assistant. That s what is missing on current rhasspy project to one day compete with the big one.

1 Like

Mycroft is small enough that I’m not worried about not having enough sway. Offline speech to text and text to speech are already being promised for the Mark II’s release :slight_smile:

A difference with Rhasspy, of course, is that we need to find some way of funding development. I proposed that we offer to train custom voices or speech to text models for businesses, and use that money to keep the lights on.

I absolutely agree (though I would argue Raven is even worse :laughing:). I’ve already started working a new wakeword system (based on this paper). Let’s hope it performs better in the end!

No, I’m not abandoning the project. I have some stuff in the works for 2.6, in fact! Besides spare time, what’s holding me up right now is that so many things have changed in the past year that need to be integrated.

For example, Larynx has grown up as a TTS system and (as “Mimic 3” under Mycroft) is fast enough to be useful on a Pi. Additionally, Vosk and Coqui STT are now mature and ready for use (though they can’t be re-trained like my Kaldi system).

Another exciting thing I’ve been working on in a hybrid STT system that can recognize fixed commands and fall back to an open system like Vosk/Coqui for everything else.

Thanks! I think the future looks bright :slight_smile:

3 Likes

What do you mean by this?
I ve been using coqui and deepspeech before this in my homebrew nodered pipeline and i train a custom scorer based on my own domain specific language model which also adds new vocabulary and i actually find it it a lot easier than kaldi for this.

I mean re-trained from scratch quickly on-device. You can definitely create a custom scorer for Coqui STT/Deepspeech, but adding new vocabulary/sentences to the pre-trained scorers isn’t possible (as far as I know) without recreating the language model or doing an expensive merge on the n-gram counts.

Both Vosk and Coqui STT let you boost existing vocabulary at runtime, which is awesome. My goal with the hybrid STT is to allow for fast re-training of fixed commands, but have it “know what it doesn’t know” and let Vosk/Coqui do what they do best (open-ended transcription).

1 Like

Dunno @synesthesiam phonetic pipeline based systems are ok, but newer ‘end-to-end’ ASR seem to be providing better accuracy nowadays even if it is a single all in model.
‘end-to-end’ aint the best description but that is how approx differentiate the 2 also ‘end2end’ also can be lexicon free and being possible to handle out-of-vocabulary (OOV) words so don’t need any re-training.
Such as https://github.com/flashlight/flashlight/tree/main/flashlight/app/asr

If only googles new tensor offline ASR was opensource but you can only wish.
Why fixed commands as are they not far more inflexible than a end2end asr with intent decoding by NLP?
I thought the infrastructure of Rhasspy was for low load whilst isn’t your target now a Pi4?

ASR would benefit from transfer learning due to dataset and model size where a local capture model can apply weights to a language model.
I think that is what Google are doing with their new ASR as supposedly it learns specific user intonation and word patterns.
I have never really concentrated on ASR as the input chain seems to have weaker parts of the pipeline on capture and initial keyword so never really progressed further up.
The quality and consistency to dataset of capture is really important and if its not right at the start further up the pipeline recognition will degrade.
So with the new audio board its an improvement even though beamforming alone is not what much state of art employs it is a step forward and guess many are wondering where this will take Rhasspy and Mycroft.

I have yet to see any end-to-end models that allow you to add new sentences/vocabulary on-device, and also run efficiently on a Pi. Oh, and don’t forget that more languages than English exist :wink:

To me, the hybrid approach (fixed commands + open local ASR) is what Rhasspy is all about: offline user-defined voice commands. The added flexibility of a fallback is great, but the overall point is to have the user train the system and not the other way around. I want to say a command in whatever way I want, pronouncing words the way I do, and never have it leave my house.

I am focused on the Pi 4 now, but with 2GB of RAM. So the flashlight models you linked (AM + LM) couldn’t even be loaded!

1 Like

Yeah guess your right Pi probably isn’t the best platform for AI lacking a decent GPU or NPU.
End-to-end if running with a lexicon then just add to lexicon, flashlight is a research framework and C++ so didn’t expect you to be running it, its just has parameters for all and was posted as an example as one.

I was opening things to discussion as the old system is getting quite dated to the rapidly moving voice AI scene and was wondering if you where working on something more current and flexible?
2gb is more than enough for many models, as far as I am aware there wasn’t any models in the link provided just some details of where Facebook research are publishing.
The old system could with a squeeze reside on the original zero and is sort of indicative, where a 2gb pi4 has considerabilly more scope even if GPU it aint great for AI acceleration and lacks a NPU.

I think its probably had its day but if you are going to take the effort to maintain the same then great.

Reason why Flashlight is of interest is that it is C++ though as its interesting as Rhasspy like infrastructure is now available on microcontrollers and I don’t agree the models are huge as like the ESP32-S3-Box demonstrates you can and is much more cost effective and has the added advantage audio processing lends itself to RTOS DSP of a microcontroller than application SoC.

I was purely wondering if you where working on anything new as this arena has been hugely fast paced and changed dramatically where industry leading offline ASR is embedded into mobile phones and simpler systems are utilising tiny low power devices such as ESP32-S3.

Rolyan how are you going with your own ESP-S3-32-Box ? Is it ready for real world use ?

The ESP33-S3 hardware does have distinct benefits for AI applications, and the demo looks impressive (as demos are supposed to). The demo appears to oversell it as capable of acting as both Voice Assistant and full-featured Home Automation controller … yet I see a big contrast between the espressif/esp-box repository and the activity on Rhasspy or Home Assistant repositories.

I don’t recall anyone ever suggesting Raspberry Pi was suited for AI. What it is, is affordable and (until recently) a freely available general purpose platform. Rubbish it all you like just because it isn’t your ideal platform, but I don’t see Raspberry Pi going away.

Look you can be a Pi fan and say it is affordable and yes ESP32-S3-Box works and because it has a audio pipeline containing AEC + BSS in tests it works better than a Pi with Rhasspy lacking simple audio processing.
Its not just Rhasspy as all the linux hobby projects have been missing essential initial audio processing that is an absolute must for what is considered basic voice AI standards.

It is not affordable as a voice AI as Mycroft clearly demonstrate with a $300 unit that offers little over a $50 unit and is completely inferior to $50 to $100 commercially available product and that is reality and yeah some hobbyists will build for fun but that is all they are doing.

There are loads of projects that the Pi does really well but the lower end of original zero to even Pi3 running Python for voice AI doesn’t work well because of specific reasons I often mention because I am being objectively honest and not just a fan boy.

You would not run a Voice Assistant and full featured Home Automation system because they are functionally distinct and benefit from running on distinct hardware that benefits them.
Cars don’t have toilets because generally its considered there are better places to take a dump and bloating a singular system is generally bad practise that often will land you in the shit.

But all the above is not the question or has anything to do with what I was asking as I was presuming because of Mycroft and as synesthesiam confirmed the focus is now a Pi4 2gb which does have far more processing power than the initial zero and asking if there is anything new in the pipeline of anything more capable than a system that had very modest roots.
With TTS we have seen this with larynx which really needs a minimum of that 2gb Pi4 64bit to really run well and all I am doing is asking if there is anything planned.

The ESP32-S3-Box was purely a demonstration that ASR models are generally getting smaller and I have no idea where synesthesiam thinks there are models that will not fit in a 2gb Pi4?
I mentioned flashlight to dodge my opinion that I feel VOSK is now a better option and other elements have evolved whilst the core ASR has pretty much stayed the same whilst elsewhere rapid changes are being made.

So how are you doing Donburch with your own Rhasspy Pi ? Is it ready for real world use ? As I don’t make any false claims about the ESP32-S3-Box as some others do with certain hardware and infrastructure.

The thread is ’ The Future of Rhasspy’ and I was asking as in certain respects it has stayed static.
It will be interesting what Upton says on the 28th The Pi Cast Celebrates 10 Years of Raspberry Pi: Episodes With LadyAda, Eben Upton, and More | Tom's Hardware as hoping we might get something like a Pi4A where the A is AI and minus USB3 the spare PCIe lane brings on board a Raspberry NPU as the PI is starting to lose huge ground in this area.
But that is just discussing the future and what are likely becoming essential requirements.

I personally feel the big processes of voiceAI TTS & STT can be shared centrally be it X86 & a GPU or what I have preordered Rock5 with 6Tops NPU that employs many ears of distributed room KWS to finally get real world use for low cost.
As if a application SoC is not a great platform for audio DSP processing then partition process to what it is great for and that is how I see Raspberry and satellite ESP32-S3 KWS and use both for what they are good for and not for what they are not.
Also waiting for the Radxa Zero2 which has a 5Tops NPU but until we get the cost effectiveness of a Pi with NPU currently unless light load the Pi is not a great platform for AI and that is just fact.

Dan Povey talk from 04:38:33 “Recent plans and near-term goals with Kaldi”
https://live.csdn.net/room/wl5875/JWqnEFNf

Tara Sainath End-to-end (E2E) models have become a new paradigm shift in the ASR community

Do you have anything to share that could be the future of Rhasspy or any cost effective VoiceAI?

Thanks for the link, it was great to hear the latest from Dan. I agree with him on the need for lexicons of some sort, and am happy that their new Kaldi stuff will stay on that path. It’s also becoming clear that I need to seriously consider using byte-pair encoding as an alternative to phonemization.

Please consider the bigger picture when making such statements. Why is that $50 unit $50 and not $300? Because they can manufacture a million of them and sell them at loss. Why? Because it’s more profitable to spy on you than to just sell a smart speaker.

I think the most important next step for Rhasspy is finding a way for others to more easily contribute updates to existing services or add new services. Changes are happening so rapidly that I obviously can’t keep up.

I’ve struggled for a while to come up with a better architecture that would allow for people to easily download new services, but there are so many unique use cases that I keep scrapping it :confused:

I have and the bigger picture is Googles next gen ASR / KW system is completely offline and is only online when you use a service as in grabbing the news, weather or play music, youtube or whatever.
There is no bigger picture when it comes to such a commercial difference that because they can subsidize product with services but much of the cost has nothing to do with selling them at a loss purely the huge economies of sale the big guys have.
Its likely we are not far off next gen smart AI with onboard NPUs like the Pixel6 giving approx 4 TOPs in 5 watts to enable offline, offgrid ASR, but maybe quite a few years yet and we will have to wait and see, but sadly no new announcements on their 10 year birthday from Raspberry.

Its interesting to watch you go into Mycroft sales speech as now I guess you have to but for me what $300+ does buy has some real cool alternative options that offer more, sound much better, look much better and work much better and many of them are less.
That is just my opinion and I am going to sit back and see how you guys do and how the reviews come in when released, but much of the cost of the MycroftII is down to the design and economies of sale chosen.
I just have a minimum level of expectation and because I have used the latest full Echo & Nest audio and have a reference and even though dev wise I tinker a Rhasspy or Mycroft would likely end up with a strop and the bin if for use.
I eventually went for 2x Nest Audio in a stereo pair that I managed to pick up for just over £100.
I don’t use them all that often but when I do its mainly music and news whilst I am doing something and those services are online anyway which for me is Spotify free and I put up with the adverts.
I think the Ech04 sound better and also have a zero latency aux in and went the Nest Audio because I think the recognition is slightly better and that is what matters to me with a voiceAI and the disparity with opensource is huge and my biggest problem and privacy becomes a tin foil concern when things run so bad.
But hey that is just me and I keep my interest up purely with developments and what is current with hardware and opensource and there is some very interesting stuff out there but for me it not Mycroft.
When my privacy is going to cost me $300 and not work to my expectations whilst I can not bother that Google & Spotify might have an inkling of my taste in music you can guess what I am going to plum for.

I am here because of my interest in AI and generally whats happening but your spy scare stories mean very little to me as I am still likely to use VoiceAI for online services.
The only thing offline is occasional alarms as they are in the kitchen/lounge of a relatively small flat and meant I could ditch the HiFi as actually for what I use they are good enough but again where the MKII is lacking.

I’m not trying to scare you, I’m just saying that spying, etc. is part of the total cost. Like with environmental externalities and poor labor practices, sometimes the final consumer price is not the only thing that matters.

I don’t have to, but I would certainly like to continue working on open source voice tech. It would be especially disappointing to have the Mark II (and Mycroft) fail because people who don’t value privacy over price go around complaining that it’s not an Echo.

Its not that its not an echo its the stupid choice that it is trying to be an echo and failing in just every area, so I might as well have an echo.

Why open source is trying to copy verbatim consumer electronics but failing when with the diversification of use is very obviously suited for client server and expense can be shared but it is not want for me its just Mycroft have chosen that path but in every aspect its inferior.
Why you have chosen a product model that is likely not efficient anyway has nothing to do with me wanting an echo but if I am going to buy something like that I am not going to buy something that is so inferior.

That is Mycroft’s problem and not mine and you can try your pitch as much as you wish but the choices they have made for me are bizarre and success and failure are in the hands of Mycroft.

By client/server, I assume you mean someone buying a server and having a number of satellites that use it?

You don’t even need satelites you just need KWS ears and a server nowadays can be a ARM board with NPU even a Pi with a Coral USB but even then a Pi starts to rapidly become less cost effective than it may 1st seem.
There is a huge amount of capable and cheap 2nd user equipment that is capable of multi threading.

As for struggling for infrastructure the problem has always been adoption of infrastructure without need mainly for branding.
A more loosely coupled modular system of feeding back to upstream projects of larger audiences has always been a better option for me and have been critical of the choices and bloat for a really long time but never bothered along that line as the start of simple low cost audio processing was always missing from the pipeline and we are beginning to see products that are finally filling that void.
That has been a huge hurdle as the very start of the audio process of voiceAI has been missing but yeah my Pi4 or old X86 refurb machine needs only a single unit to act as a ‘server’ and a room can have initial audio processing done on a $20 microcontroller of a KWS ear for each room.
So yeah a NUC, SBC or even your old desktop or laptop can be the basis of a central server at extremely low cost and cover numerous rooms.
Voice control and capture is extremely scalable because of its manner where short singular infrequent commands are often the norm.

Its very possible to add a coral accelerator to a PI 4 and broadcast audio to a wireless audio system than embed that functionality purely to call branding of your own.
Its likely you could do a better system for 3 rooms for half the price of a MKII that comes close to $1000 if you want x3 MKIIs.

And no not satellites as I have always argued they are bloat and want to get away from that bloated term, as all that is needed are network KWS mics (Networked Ears) and a single station.

1 Like

Before I comment on this topic I’d like to say a belated congratulations to @synesthesiam on your new job and I hope you can contribute to making Mycroft a real privacy based alternative to the other commercial offerings.

This has been an interesting discussion and you both have good points so I thought I might give another opinion which I think overlaps both points of view.

Personally I value privacy over price but only to a certain degree.
When they came out I bought 10 echos, of which I now use one as an alarm clock and timer, one other as a timer and one for general questions, the others basically are not used any more, largely for privacy reasons.
I have been struggling with Rhasspy for the last 2 years now because, while it is an excellent product for setup and flexibility, at least in my environment, I just can’t get reliable kws and accurate voice recognition at a reasonable price due to the quality of the microphones and the cost of building out satellites that then are fine in a close quiet environment and become almost useless when you try to use them in real world conditions.
I am running the base on a nuc with TTS also running there and mainly want to be able to control my home automation and music playback to sonos speakers from my local library, all which I can do now in my test lab but I would not put it in other rooms for the reasons I mentioned above.

Also I have been watching Mycroft since the first model and have been reluctant to go near it due to cost, capability, reliance on the cloud and it just doesn’t look professional.

I agree with @rolyan_trauts in what I would like to see:

  • Central processing/server based. If this was around the current price of the MarkII I would gladly pay. I would even pay a bit more if it could do the rest.

  • Low cost modules for microphones/satellites I could deploy to each room. By low cost I mean comparable to the echos and nets of the world or a little more (the privacy and flexibility would be worth it)

  • An easy way to integrate my own data and systems into it (i.e. slots)

  • A standard interface or protocol api/websocket/mqtt (pick one or more) that would allow things like node-red/home assistant or any other system I decide to build to integrate with it for I/O and control

I like almost everything about rhasspy except the satellites, I agree with @rolyan_trauts they are too bloated and the hardware they support is expensive for what should be needed and I agree with @synesthesiam about paying for my privacy.

If Mycroft could offer something like a base unit for the heavy lifting and packs of “microphone” units for the user to place in rooms I think they could offer a much more cost effective solution and using my own example I would seriously consider paying $1000-1500 for a base and 10 microphone units where I wouldn’t buy a single Mycroft Mark II for $300. I would even stretch my budget higher if enough features were offered on the Mycroft platform.

In the mean time I have echos for the menial tasks and keep trying to work out how to get rhasspy working to a level I would consider presenting to my wife rather than inflicting on her.

2 Likes

The best bet is probably the esp32-S3 but we are going to see a load of very capable low cost microcontrollers that are perfect for wireless KWS roles.
I am still have made no progress as they have released another product called a esp32-s3-box-lite and I really don’t care about screens, output, speaker its purely the input to be able to run a descent KWS and have a speech enhancement pipeline.
The original esp32-s3-box had a analogue loopback from the dac to sync the AEC ref on a 3rd ADC channel so the idea of just putting 2x I2S mics on a standard dev kit became a show stopper.
The new lite box has a 2 channel ADC so hopefully there is a software update on the way but also supposedly you can just buy the board alone.

I am going to keep calling them ‘KWS Ears’ to stress how simple needs are but why I keep focussing on cost is not just the cost of a singular unit for a room its that a room could contain multiples to form a distributed array microphone.
A softmax probability from a single KWS is a good enough metric if the highest value in an array to use that mic for the current ASR sentence. The ones that didn’t hear the KW are likely not hear the ASR sentence and its that simple as each ‘ear’ is completely ignorant of another’s existence.

Hopefully the ESP32-S3 boards will follow the same economies of sale that previous ones did and maybe get as cheap as $5 as you could have 2 or 3 in each room if you so wanted to as each additional mic can be placed to provide further isolation from noise sources and making far sources nearer.
I am not on a Espressif sales pitch but its the only source of free speech enhancement I know of AEC & BSS I don’t like how the base libs are blobs but hey, if another microcontroller comes along then hey but Espressif does have a history of being extremely cost effective wireless microcontrollers.

I don’t want the ‘KWS ears’ to be a Rhasspy, Mycroft, Sepia, Project Alice or one of a plethora of projects all doing the same I just want to set a basic ‘KWS ear’ system that is simple interoperable with all and doesn’t pander to any other projects protocols.

I couldn’t care less about branding or ownership its just a very simple websocket client/server queue that is file based on zones and acts as a bridge to the input of any ASR.
That is the only dictate as the zone file structure on the input matches the output to where ASR txt is dropped.

Audio is coupled by a linux asound loopback not some weird and wonderful protocol other than the websocket of the server side of the KWS bridge. Probably doesn’t even need the file system as likely the current sink of a loopback is more than enough info.
Its why I have been waiting for product as it will be built up from the audio source path with thought to always boil down to the lowest common denominator of simplicity and interoperability without bloat.

Its why I have stayed on the forum as there are many actors here @synesthesiam to say the least :slight_smile: but I find it infuriating that many projects are applying their own methods purely to have ‘their’ own methods even though the Sepia initiative seem to be trying to address this.
This is Linux this is opensource and all the pipeline stages of VoiceAI are distinct and we should be able to partition and give choice to any as Linux and opensource does.

I am sort of critical to the MkII of Mycroft but what they have is an excellent skill server and wish they would concentrate on that as yeah I would love to tack that onto the end of my ASR of choice and so on.

But going back to ‘KWS ears’ I think opensource can do more, be better and more cost effective but its sheer stupidity to try and verbatim copy commercial offerings as likely your going to fail and there could be much better ways of doing things, for less and one of them is integration and reuse and interoperability.

I haven’t ruled the Pi either as Zero2 & Pi3A+ are both great products but until someone does provide effective AudioDSP utils I have a bottleneck for even the base function of an ‘ear’ its still a platform that easily installs network synced multiroom audio such as Airplay, Snapcast & I think squeezelite (is it synced?) That will take considerable work to port to microcontrollers where efforts with the original ESP32-S3 where slightly too constrained.
The 2mic & 4mic hat on the Pi are extremely cost effective and all is needed is the efficient code we run the rest of out Linux audio system on which isn’t python and its sort of sad as the hardware is capable.

Squeezelite can play multiroom usin Logitech Media Server / LMS and there is also a project using ESP32.

There is also a complete device with Stereo speaker, Microphone and port f.e. a display.

1 Like

Yeah I have been thinking for some time that audio out of a Linux AI would be interoperable and about choice.
There is no reason to brand and embed audio out into a opensource voice ai as all is needed is an opensource network synced audio player.
I know Airplay & Snapcast use NTP to sync latency of network sinks in the same room and also with rooms. It doesn’t remove latency but just ensures the audio out is in sync to various clients over any network.

Snapcast is full blown opensource but is a tight fit into a ESP32 also I think someone has ported Airplay1 and I think Squeezelite is ported but I don’t know if squeezelite does NTP audio sync?

To use opensource and feed back up source and become part of the herd strengthens opensource whilst embedding slight changes into your own system could be considered leaching as it dilutes a herd and becomes harder to maintain via a smaller herd as its deliberately been made proprietary.

Esp32 has a strong analogy to the Pi3&4 as ESP32 is based on a LX6 processor architecture and the newer ESP32-S3 is based on LX7 architecture and can have more flash, ram and is over 10x faster than the earlier base model with various operations.
So even though a ESP32 is a tight squeeze that is a base model with no additional PS ram but the extra oomf of the esp32-s3 and that flash, ps-ram and internal ram has all increased and being new its options are very open.

I think its very possible to run KWS websockets & a network audio system on the same microcontroller with the advent of a S3.
Also its not necessary to embed audio into weak insubstantial speakers as it can be much more cost effective and superior audio wise to share a wireless or wired audio system as network ears only need to be ears and audio out can be from a room station, central station or a network speaker, even hdmi cec or IR can wake your TV and use its speakers.

Reuse and interoperability can be a superior system that is often more cost effective.

If anyone knows if squeezelite is network synced please say as out of the 3 its the one I don’t know that much about.

Google & Amazon with various cast in inbuilt sync functions are embedding deliberately to call there own for branding ownership and needs to be dodged.
Strangely with apple Airplay is more open.

https://forums.slimdevices.com/showthread.php?112697-ANNOUNCE-Squeezelite-ESP32-(dedicated-thread)

I say open but cracked which I am not to sure about the Google cast system.

airplay1 definitely is being used.

Also are there anymore opensource audio sync projects with good support herds?

On a Zero2, Pi3 or 4 very easy to implement and run.

Does anyone know with Pi HDMI cec if you can change source to and from the Pi.
I am pretty sure you can change to a hdmi source not so sure if cec on one hdmi can switch to a source of another.
Think its time to ditch my DVI monitor and start having a play.

I’m a bit late to the party, but I though I’ll chime in as well.

For me the most important features of Rhasspy are:

  • Completely offline operation of course. In fact I’m using it in places that don’t even have internet, I used my phone to download the models.
  • Guided setup. The web interface is awesome. A config file would be fine as well, but the important part for me is that I don’t need to search through the forum to find out what options there are and where to find the current models and such.
  • Good recognition and acceptable text-to-speech quality. She (the default Larynx voice is female) does get it wrong sometimes when there’s noise, especially numbers (“set the timer to 5 minutes” often gets me 45 minutes) but in general it’s pretty good, especially for my very basic microphone setup. TTS ist also quite good, especially compared to alternatives like Mimic. (She does have a few impediments like she pronounces 14 as “fourtheen” or “partly” as “part-lie”.) I just wish Porcupine was a bit more robust, but I want to switch to satellites, so I’ll have to look into wake word detection anyway.
  • Speaking of which: Satellites
  • The ability to control pretty much the whole flow via MQTT. Dipping the music when the wake word is detected, cancelling the recognition with a button, etc. This could actually be improved, like triggering TTS or playing a sound file via MQTT. Or even switch between different training sets, that would open up the way to have an actual dialog. (“set the timer” - “how many minutes” (switch to number set) - “23”)
4 Likes

Maybe: https://www.ngi.eu/

1 Like

I think Google are already far in front with this Google Seeks Help From People With Speech Issues - Disability Scoop

They have been working with disabled people with Project Euphonia for several years now, don’t know what product they do, but like Rhasspy in comparison to Googles new offline ASR there isn’t really in terms of results.
Disability software needs to be effective, low cost and accessible and not sure that it is opensource is a criteria of any importance and its near impossible for a singular developer to collate the datasets they have so you will always be inferior.
Its one thing Rhasspy & Mycroft have seemed unable to garner much and that is datasets Mycroft had a lot of web site blurb at one time but never anything as a dataset not even ‘Hey Mycroft’ seems to be available but might be because I am not a fan of Precise I never looked hard enough.

I dunno what they have planned but sometimes Big Data can be beneficial, opensource needs the datasets as until they have something like what bigdata has we can never compete.

For the disabled if it entraps or not to an Android ecosphere android.speech  |  Android Developers is prob not much of a consideration but still have to see how they release there new tensor chips that are basically ultra low power npus and its really hard to compete if they release there Pixel offline self learning ASR for disability as its really hard to recommend what is currently opensource to the disabled.

Hey @synesthesiam ,

I know this thread is a few months old now but I wanted to say congratulations as well. If you are happy then thats all that matters

In reading some of the posts here, I see a lot of people have issues with MyCroft’s marketing of being privacy oriented but then there practices being the opposite. I share these issues but I just wanted to point out the fact that they hired you. Someone they know will be pushing for staying open source and going completely offline. So although im not going to start using MyCroft yet, hiring you seems like a very promising sign that they will move in the right direction.

I have been using Rhasspy for a few years now and one of the things I love about Rhasspy is the ability to customize. I know that normally comes with being open source but with every component of Rhasspy, you have the option to use a local command. Which means if I want to use a TTS or STT or anything that is not integrated out of the box with Rhasspy, I can set it up myself with the time and knowledge. A skill store and dedicated hardware option would be great for ease of use and I am all for it. But as long as we do not lose the ability to customize in exchange for ease of use. So i guess what im saying is please make sure MyCroft stays open source haha, and the options to customize any component is available always.

I believe you mentioned that other companies you interviewed with required that you shutdown Rhasspy if they hired you. Honestly that makes sense to me as Rhasspy would be competition to the company you are working for. Seems like a conflict of interest. So that lead me to wonder why MyCroft didnt require the same thing. My hope is that its because they have seen what you have done with Rhasspy and truly do want to merge the two together.

If you say you are not abandoning the Rhasspy project then I whole heartily believe you. You have never given us a reason to doubt it. If there is anything this community can do to help keep this alive, there are a lot of people, including myself, that would be willing to help.

Thank you for all your hard work thus far and all that is to come

1 Like

Congrats on the job!

Now let me say that I tried Mycroft before Rhasspy a few months ago and it was fun, but not nearly as easy to customize as rhasspy. For someone who has trouble saying certain sounds (particularly explosive sounds at the beginning of a word), the option to choose a wake work and set custom commands is amazing. It’s also easier to get it to do exactly what I want in Home Assistant using intents and writing custom services in pyscript on Home Assistant than it was to use the Home Assistant skill in the MyCroft store.

For example, my wake word is “grasshopper” and the command to turn on my lights “lumos” and off is “nox” which are both more fun and much easier than saying “turn on/off/out my lights.” Oh yes, I also set “out” to also mean “off”. I suppose I could write a MyCroft skill to do all that, but this is so much simpler.

So, please, keep rhasspy going or integrate that type of functionality into MyCroft if you can. It’s just such a good feature.

1 Like

@dblanc28 Thank you! I’m definitely happier at Mycroft than I was at my previous job :slight_smile:

Yeah, and as @VoxAbsurdis pointed out, their privacy policy is currently awful. Fortunately, the (now) CEO Michael Lewis is having the Mycroft legal team rewrite it from scratch. Going forward, the default is going to be not storing any user data, where possible. Cloud-based STT and TTS are going to require some additional thought, of course.

This is a central feature of the new architecture I’m designing for Mycroft 2.0 :slight_smile:
The idea is to have a small supervisor process that runs each service as a regular program, and passes events through their stdin/stdout pipes.

With some minor configuration, this lets you reuse or adapt existing programs with minimal effort. For example, arecord and espeak-ng can be used as-is for mic/TTS “services”. Your STT could be a curl call, and your intent recognizer could be grep :laughing:

This is my impression too.

Hi @kicker10bog, thanks for the feedback!

I’m pushing to include this as an easy option. To me it seems like such a no-brainer; rather than struggling to decide which sentences should trigger which intent for everyone, you can do something crazy: let the user decide :exploding_head:

The only thing I would consider changing is how sentences/intents are specified. It’d be nice if the format was portable between open source voice assistants, so people don’t to rewrite if they find something better. But I haven’t found other formats that have as much flexibility as my format :confused:

1 Like

I’m late to this conversation but I wanted to chime in. I had Picroft set up for a while, and started working on skill development, but I lost interest quick due to some other issues. I’m a supporter of the Mycroft Community, have donated as well as have the subscription. I had preordered the Mark 2 but some things have kept me from being able to grab it (or the dev kit when it released). I’m keeping an eye out for when they release the custom board that can be used with a DIY setup!

My main problem is the online dependence. It’s not a privacy thing for me though. I want all my smart devices to remain controllable by voice even if my Internet is down. Alexa only does it for Zigbee devices, and only if you use her as the gateway. You can probably figure already I have a universal gateway (going to upgrade it soon tho for full Matter compatibility) so that doesn’t work. I hate having to bring up my HA app every time I want to control certain devices, and I’m a bit miffed that my control panel tablet is overheating and shutting itself off after about 12 hours.

It’s good to hear both projects will continue. I haven’t tried Rhasspy yet, I’ve been meaning to which is how I ended up here today. Planned to install it today :yum: I haven’t tried the Mycroft self-hosted back end yet either. One thing that I’m sure is not ready yet is either using SMB, DLNA, or n emby/jellyfin/airsonic/etc plugin to create a json formatted database of your media collection and be able to call up my own music to play on my Yamaha MusiCast receiver (over DLNA). I run an unRaid server and have 5 echo devices in my small 2 bedroom condo :yum: also 2 Google nests, a hub and mini. I keep meaning to switch over to Google because Alexa is very limited and there’s still no multi-commamd mode.

Will you basically be porting Rhasspy to combine the best of Rhasspy and the self hosted backend for Mycroft?

I’m still debating how I want to do this. I have a Pi4b 4gb ram collecting dust and a mini PC with a Celeron j4125/8gb ram just running OSSIM currently, but might move that over to my unRAID server, freeing up the extra mini PC. I have a ReSpeaker 2 board and Logitech speaker still from PiCroft, but those will work with the mini PC as well. I’ll need a couple satellites too. Pi Zero W or Zero2 W should be fine for that, no? I also want to finish integrating Tuya into my nodemcu projects scattered around the house :crazy_face:

Anyway a late congrats on the job, think I’ll pick back up on skill development too.

1 Like

I just plain don’t trust manufacturers to not defeature my things or even something as basic as to stay in business.

I don’t buy devices that are not standalone.

You want me to buy your doohickey?

I have to be able to use it without Internet.

Offline voice recognition is why I am in your bubble… or vice versa.