Is there no more pi microphone array hat?

Matrix Voice is completly dead and for Respeaker there are only relealy old driver from HinTak, that don´t work with the newest kernel. Are there no more alternatives?
Will i have to use an usb mic array? Whitch one works best (with rhasspy) 64 Bit OS & Kernel 6.x and get future updates? Read something about Anker PowerConf S330, but one the website there are no drivers for linux!? Is there no more hardware with official linux support?

I am confussed.

My matrix voice works fine,but on a very old system. I now want go give rhasspy 3 a try and don´t want to start with old Kernel and 32 bit. :frowning:

Thanks a lot, Mili

2 Likes

Come on … you all have to use a mic! :wink:

Come on … have you read the previous conversations ?

To repeat once again … if you wast Rhasspy 3 with Home Assistant - that is called HA’s Voice Assistant and support (including setup tutorials) are over in the HA community forum.

I note that HinTak updated his drivers very shortly after the latest RasPi OS was released. Have you tried it ?

My latest satellite (on my test bench) uses a USB mic and regular speaker, and seems to do quite OK without the expense and driver issues (ecxepting that passing USB to Docker can sometimes be a pain).

1 Like

Thanks for your answer. Witch previous conversaaion do you mean?

HinTak´s last update of the driver is 1 year old. I have try it and get this issue:

This issue is more than 3 mounts old and nothing happen. So I don´t know if it ever will be fixed.

I don´t use Home Assistant and only want to take a look on the version 3.
But also I want a new system for my normal rhasspy, with the newest software.
And HinTak´s work is nice, but what will bring the future? Is there no more company for mic array - where i hope i get a little bit more time of support?

And you wrote you use an USB mic, but not which on? Also the docs are so old and only Matrix and Respeaker are shown there. Is there no list of good supported hardware? A normal mic is no alternativ to me, my Matrix voice works in range of 4-5m, and i don´t want a downgrade.

I have taken to using usb speakerphones for my assistants, one of the main benefits being that they have hardware echo cancellation built in so your assistant won’t hear itself talking. If you use your device to also play music, it can filter that out also so it can still hear you. The drawback is that the sound can be tinny, especially on cheaper models, but all the ones I’ve used have had far-field microphones that work well enough for speech to text. The tinny sound only bothers me when I’m playing music. You can find them for under $20, although you can certainly spend a lot more.

Hintak has previously commented that he is only updating for newer kernel versions - not doing any updates to functionality - yet I have always been surprised that HinTak just copied seeed’s readme including the links to the old OS and seeed’s forum (where they stopped responding to users 5 years ago).

Seeed’s demo includes software features not included in the device driver, and it seems that all companies are viewing these voice processing functions as proprietary :frowning:

As for the reSpeaker recent update, I had simply noticed

install.sh     Fix /boot/config.txt for bookworm installs     2 months ago

And I have no clue about the kernel panic. Lack of ongoing support is a main reason i don’t recommend the reSpeaker HATs.

You are totally correct that the info on rhasspy website and community is old, and very few updates for about 2 years since Mike started working on rhasspy 3 then got the Nabu Casa job. I am hoping that now HA Voice Assist is established he will be able to spend some time documenting generic Rhasspy 3. I suspect that all the code/components are available - they just need more detailed documentation for people to see how they fit together.

As for currently recommended hardware …

  • My generic USB desk mic + speakers are fine while sitting at my desk half metre away.
  • raspi + reSpeaker is rather expensive for what it actually does.
  • Some conferencing speakers have the extra smarts we want, but too expensive to deploy all around the house - on a pension in Australia anyway
  • Home Assistant Voice Assist are currently recommending the Espressif ESP32-S3-BOX device for voices satellite … though they are using Wyoming protocol.
  • I personally think Nabu Casa could start with the ESP32-S3-BOX and improve the speaker, mic and case acoustic design to bring it closer to alexa/google sound quality.
  • Hopefully @rolyan_trauts has been putting his apparent expertise to use and fill the market gap he sees for good quality affordable off-the-shelf “ears”
2 Likes

Thanks for the detailed answer! :slight_smile:

@ aaronc - Which one do you use exactly? How many meters can you go away from it?

@ donburch - Do you know some conferencing speakers thats works fine? I don´t need so many of them, an therefor the price is not so important for me.

PS. I searched for your notice to the reSpeaker update and with that i found:

Befor I only have found:

I will give the x version a try, thanks a lot for this help :slight_smile:

That bit is quite easy and doable for a while but for some reason upstream uses Whisper which that has been trained on a specific type of audio input of monaural near field recording with tolerabilly low level noise.
You can fine tune Whisper but its a very large model and much work and still likely worse than a much smaller trained for model.

I keep repeating that the big guys have a hardware advantage because they dictate the ‘ears’ so the models in the cloud are trained on data processed by them.
You can feed Whisper various great filters and speech enhancement algs and WER actually goes up as the signature those algs create where not in the training dataset.

I never suggested the ESP32-S3-Box as for ears it has much expense that is just bloat, but dev on a just esp32-s3 and mics only, never took off.
There are various projects but really the esp32-s3 is just being used as a wireless mic and has zero pre-processing but just feeds whisper.
As usual I am bemused especially the one that cracks open Google nests rips out a working audio preprocessing system and just sticks in a esp32 as a wireless mic.

If you could get upstream just to commit then likely the answer is yes, but you need to preprocess the dataset and train a ASR to the audios input signature, for greatest accuracy.
When you do that, things work really well, like commercially units have been doing for a number of years.
The bring your own to the party means we have what we have with just single mics feeding Whisper as no-one will dedicate a ASR to specific hardware that has a certain signature.

People are just having a bit of fun with Year of the voice and the projects that you can make, but there is nothing remotely near the dictate you need to even come close to the big guys smart speakers.

Its very simple upstream ASR trained on data preprocessed with the filters and algs of hardware of use, provide masively better results than the quantum task of trying models that can cope with, bring your own audio to the party.
Whisper does an excellent job of this and why its being used even if custom designed solutions of hardware dictate aka Nest, Alexa, Siri units in use work vastly better.

Plugable do a USB stereo soundcard as do Axagon

Respeaker still do the 2 mic.

All else really are totally pointless because we don’t really have any algs to run on them.
I did do a simple delay & sum beamformer for any stereo mic input that because of the simple summing, likely has a very small signature and is similar to single mic and would work well with Whisper (You need to test with Whisper as it may actually not like the speech enhancement of some devices)

Likely a USB soundcard with AGC and mic is as good as the latest and greatest due to how the Whisper model reacts to the audio input.

There is some nice new hardware about.
Radxa Zero3W which OKdo ROCK ZERO 3W 1GB with Wi-Fi/BLE without GPIO - OKdo is an £18 SBC that for ML is faster than a Pi4.
Also the ultra cheap Radxa S0 a £11 ultra low wattage A35 SBC
OKdo Radxa ROCK S0 512MB Single Board Computer Rockchip RK3308GB with Wi-Fi 4 / BT5 - OKdo
Strangely both have onboard audio codecs but don’t think they are on gpio … ! ?

Hi Stuart,

I thought this conversation was about the “ears” - getting high quality audio from the distributed devices - which I expect should be quite separate from the Speech-to-Text phase.

I understand that the ESP32-S3-BOX was designed to demo a variety of things that the S3 can be used for, and so is overkill for our needs - but I’m sure it was you who pointed out the extra instructions in the ESP32-S3 processor should allow more of the audio pre-processing functions (which you are intimately familiar with, and I just refer to as magic) to be done cheaply on-chip. I think that with microwakeword it’s the best off-the-shelf option at the moment.

I was also amazed that anyone would rip all the magic out of a device to replace it with a simple processor. I had expected them to use much of the circuitry on the board by disabling the existing CPU and add their own. I appreciate the audio engineering that must have gone into designing the case to channel the sound from speaker, and to the microphones, to reduce interference - and again to replace it with poorly positioned microphones :cry:

Seeed developed the reSpeaker 2-mic HAT, and several clones hit the market … and it would have been the standard by now if seeed hadn’t lost interest in supporting its driver/firmware … but seeed saw it only as a demo to encourage other companies to incorporate and support. The demo showed it has the capability, and it could probably take off again if someone was motivated to build some of that pre-processing magic into its firmware or driver. But RasPi + reSpeaker + case will always be fairly expensive compared to on-chip magic :frowning_face:

There is definitely a need in the market for a good quality, relatively cheap, satellite device; and I feel confident that it will sell well … which in turn makes it worthwhile for others to generate the upstream models for Whisper, etc. I believe you have the expertise to develop such a device, though I assume you see marketing as being the big challenge (that is my big fear). Certainly I believe Nabu Casa could build an ESP32-S3 based voice satellite device, and they have existing market presence and experience with the HA blue/green/whatever boxes).

In the meantime people can pay for proprietary conferencing mics with the magic built-in; or put up with a standard USB microphone doing as good a job :frowning_face:

I won’t respond to your comments about Whisper because i’m not sure what point you are trying to make. That Whisper is bad ? That the variety of audio input sources makes Whisper less effective ?

Nope unfortunately and it was also a surprise but models such as Whisper are trained to cope with a modicum of noise and room impulse reververberation.

Its sort of set to a near field mic of a couple of prob a max of a meter and the RIR’s that produces and relatively low level noise.
The noise levels are in the opensource docs of Whisper but also other pretrained ASR exhibit similar.
So you emply a farfield mic to strip RIRs and Noise that sound completely perfect to human ears depending on algs used can provide seriously bad results as the close field mic and moderate noise has been trained into the whisper dataset.

The esp32-S3 has some esspressif blobs for source-seperation & aec that could be used on any esp32-s32 coupled to mics with a good analogue AGC.
There is nothing special in the circuitry of the emptied Nest products but its just being used for the shell and for what is just a single channel network mic.

The 2 mic hat due to being a hat dictates how and where you place the mics and is a terribly inflexable design that makes position and insulation of your mics a near impossible chore. Most people have the hat which is a 2 mic broadside array pointing at the ceiling because it is a hat and designed that way.
Best way is to wall mount then at least the mics are pointing your way as casing will provide rear rejection.

People might not be able to pay for proprietory conference mics with magic built in as with ultra hi-tech devices such as AnkerWork S600 Multifunctional Speakerphone with Voiceprint Recognition

Whisper is trained to expect room inpulses of an open mic and the noises that may incur.
Feed it cleaned signals stripped of RIRs and Word Error Rate can climb dramtically.
No Whisper isn’t bad as it was designed for a open standard mic senario, changing that scenario and expecting it to work as well means your knowledge of what you are using is bad.

Nabu Casu likely have market presence, but have a hunch similar to Mycroft where they present a knowledgeable front.
That they are not allowing an opt-in to collate usage data from collaborating users is a massive red flag as they are ignoring dataset gold.
The custom wakeword is another red flag as even big data finds them inferior to models they can dictate and control that return vastly better results.

From: Paulus Schoutsen paulus@paulusschoutsen.nl
Sent: 30 October 2023 20:31
To: Stuart Naylor stuartiannaylor@outlook.com
Cc: Michael Hansen hansen.mike@gmail.com
Subject: Re: Just read as very short but what has been done or being said makes no sense?

Hey Stuart,

Adding Mike to the CC.

I’ve been reading along with your posts for a while now both in the Rhasspy and the Home Assistant community forums. Although you might intend well, your posts are derailing threads and you need to stop posting. Mike agrees with this.

We are not interested in hearing that everything we did is crap, everything we are working on is crap, or hear about solutions that are unrealistic and unachievable.

Consider this a courtesy warning that these posts need to stop.

Paulus

I haven’t being saying everything is crap, but everything that has been implemented is fundementally wrong.
The solutions are realistic and achievable and have been available for quite some time and are being done, just not in HA.

I have been warned not to speak about it, so I will not.
I replied because you asked, but there is little else I can say.

When I do reply as I did above and get emails such as I have posted.

Wow…

Please Stuart, take my comments with a grain of salt. They are simply from my personal point of view, trying to appreciate your point of view.

Regarding Paulus’ email first …

Paulus and Mike, sorry if I have “stirred up the Hornets nest” with my post.

Stuart I have long wondered whether you have been deliberately trolling this group to create argument - but have decided it much more likely that you just have difficulty expressing yourself, between your obvious passion for the topic, and a brain that seems to be always jumping from one idea to another. I also find communicating with people difficult, and have recently discovered that can be a symptom of neuro-divergence (eg autism). Unfortunately many of your previous posts came across as excessively critical and/or confused off-topic rambling … but after I’ve cooled down a couple of days and re-read, I have often found gems of knowledge and new perspectives.

You obviously have a lot of passion and expertise in the audio engineering field, and I suggest that focussing on what you can do right (I’m guessing the “ears” niche of the market) would be productive and satisfying for you.


Over the years i have come to appreciate that different is not always “wrong”, and that there should be enough space in our universe for different points of view.

I do understand that having better quality audio used to train Whisper (well, any STT) and feeding it better quality audio should make the whole STT job easier, and of course google/amazon have the money and marketing to get that large base of higher quality audio from a small small controlled selection of input devices.

On the other hand, Open Source STTs have to make do with the audio data that is available. Such is life. At least by promoting a cheap quality ears device the proportion of quality audio samples should grow.

Are you suggesting that HA needs to collect all the audio that users are generating, so we can build a dataset to compete with big business ? That seems very much opposite to the local control principle. Or are you suggesting that users could op-in to providing this data ? If so, how many TB of data, and wouldn’t too much of it to be low quality because of the wide assortment of mics etc. ?

And this is the thing that puzzles me most…
If HA is going in the wrong direction, and not open to listening … wouldn’t it be more satisfying to forget HA and instead put your time and effort into the other better solutions ?

Its OK as stopped posting and having opinion, I have MS with scar tissue from several exascerbations. My memory and concentration is not what it used to be.
I was using a Pi and Voice tech purely as cognative help.

All models work to the parameters of the dataset and why Whisper works so well with near field single mics and moderate noise is likely because the dataset contained such.
Its a huge unweildly model to start training yourself, but the problem will always turn up with the bring your own to the party that may have sigantures outside a models dataset.

Voice tech is a system and a pipeline all working optimally with each other, not just random opensource losely licenced that you can package and brand.

There is so much that is likely a cull-de-sac as getting the datasets for certain hardware is as easy as running the existing datasets through thoose filters and algs.
This is where the blobs of esspressif get less attractive as there is no opensource to port the code to so on a powerfull computer you can quickly much faster than realtime create hardware and filter alg preprocessed datasets by running them through on a faster computer.
So unless someone can provide opensource AEC/BSS/KWS to esp32-s3 its likely better going victorian engineering and using application SBC’s with more power and use algs we have that can scale up for dataset making. They already exist and I have been repeatedly saying so but they need to be put together as a system and if not dictate strong community guidelines on what to use as hardware.

Also the ears can absorb cost by being multifunction and sugggest Snapcast as a wire audio system and maybe even Kodi that one can plug into a TV and provide more than just ears.

Keep the ASR and LLM central as a brain or a collection of containers routed to.
I have mentioned this many times for some time, but hey.

Maybe someone can try one of these AnkerWork S600 Multifunctional Speakerphone with Voiceprint Recognition with Whisper as if WER goes up then there are big problems recommending to a comunity to buy expensive equipment that works no better and in some cases worse than ‘just a microphone’

Likely the USB chips in the soundcards I often name drop could create a HA branded mic by just adding 2x uni-directional with decent analogue AGC in a package spaced for 48Khz approx 65mm (from memory) as usb leads can be quite long and unlike a hat you can have several of them…
We have had simple solutions for a long time but they are generally ignored…
Anyone can make one but ready assembled to plug in seems to be what people want.

I lost interest after that email, really, so no bother.
I have occasionally replied to existing threads.

Really the KWS should have a preffered KW and a function to capture positive KW. Allow community to submit that dataset gold to HA as opensource with metadata.
Same with Command sentences as from pattern of use its pretty obvious which ones are positive and capture and allow an opt-in to send more dataset gold.

Everything is simple but some real basics such as database collection have been missing even though its been said again and again.

I have no idea why all this has not bring brought together as an optimised system than just what opensource they can package with Python…
It is fairly simple but from what I am seeing the oppiste is true and what is being done, is purely because they can…
Nothing is unrealistic and unachievable, its purely because they can not…

I use a Jabra 410 USB speakerphone. They are about $30 on eBay. It is plug and play and nothing special to set up if using Docker. You will need to reboot after you plug it in the first time. You may also have to modify a Linux txt file that is noted on another post on this site. Its shape is worth while as you can 3d print a cylindrical enclosure to hold a pi zero 2 and set the Jabra on top. I see a lot of comments about cost. I do not want companies listening in on my conversations. That’s why Rhasspy is a great alternative. I like cheap too, but what is your privacy worth? There may be cheaper available speaker options, but the Jabra works and I can spend more time developing my custom skill programs than trying to figure out some cryptic Linux setting for a device that I have no detailed documentation on. Off the soap box……….

1 Like

Failed to mention, my setup is using aplay and arecord for audio recording and playing options.

I am an audio snob and like my Anker Powerconf the speaker on my speakerphone doesn’t really cut it, for a smart speaker.
I am one of those that listens to a Nest mini or Echo dot and can only think ‘Urgh!’
The standard Nest & Echo are a minimum for me as ‘Play some music’ is one of the rare occasions I do use a smart speaker.

We do have 2 excellent wireless audio apps Squeezelite as it will squeeze into a ESP32 or my own fave the almost limitless configurable Snapcast.
Mounting a small amp board onto the back of a bookshelf speaker can quickly surpass Nest and Echo devices.

By using Squeezelite/Snapcast you get great zonal audio and a platform that is pretty easy to add a microphone.
I hacked together a 2 channel delay-sum beamformer GitHub - StuartIanNaylor/2ch_delay_sum: 2 channel delay sum beamformer that will extend far-field, but there is no signal to focus the beamformer on.
So it acts like a conference speakerphone and will beam to what ever is the predominant noise and be contantly shifting focus.
The methods of targetting a voice or focussing a beam is totally absent from the opensource we have even though there is code available and solutions.

Speakerphones do have AEC but we also have opensource for that even if not implemented.
With a bit of lateral thought can provide huge improvements and not need AEC by not mounting your mic in the same enclosure than your speaker.

A wired microphone can be small and very descrete even a dual mic and you might have more than just one in a room to give much better coverage.
Initially I expected ‘Ears’ to be 2x Mics on a ESP32-S32 in a small flat panels approx 75mm in width that clips onto a wall or stands on a desk.
It still needs to be powered and may contain a pixel bar/ring and even have audio out.
That is where I don’t see much prob with wired mics either as even wireless network mics still need a PSU, so really little difference or advantage.

I think laterial thinking on what a open source smart-speaker is needs some thought as with things like Squeezelite or Snapcast opensource can give Sonos like zonal audio that has a huge array of choice.
That zonal audio and where HomeAssistant can excel and beat commercial systems on functionality, quality and choice is a huge selling point that is likely under implemented and undersold.

For some reason whilst trying to escape Big Data ‘Smart Speakers’ opensource has been blinkered and tries to copy consumer individual product verbatum.
I don’t even think we need a ‘Smart Speaker’ just a zonal opensource microphone system that takes apart the commercial notion of a smartspeaker to working components of wireless audio, microphone and pixel indicators to give choice.
If someone wants to build that in a box, they can, but a modern room could be a very different scenario with a single large pixel indicator, wireless room audio and several dispersed microphones.

I guess its where you want to go but HA in terms smart controls and dashboards represents some of the most cutting edge Smart Home tech whilst USB speakerphones hanging out of mini computers is not much above Raspberry Pi Google AIY Voice Kit…

There is also other commercial equipment that can also fuse function to become more cost effective as a device as my Anker C300 dual mic webcam is great for audio and its far-field isn’t bad.
It can provide Frigate video and be a wireless room mic array and used in conjuntion with a wireless audio system.

I also tried some of the bi-directional audio pan-tilt cams you can get but found on the ones I got the mic audio is pretty awful.

You can still use the Respeaker 2mic that eventually always catches up to the latest release. Likely though again some laterial thought that USB devices are far more compatible and convient than Pi hats limited application.

There still is Pi microphone hats but really only the 2mic is of any use, but there is a whole range of USB devices that work on many platforms that prob could do with a HA mic for those who don’t want to build one.

The hardware, software all exist its just not implemented as a system and so ends up excluded.

Hi Stuart, sorry for that but could you be more concise and make your answer more understandable, and only answer the question ? If you want to share your works and your visions with everybody, it’s fine but I suggest you to open a specific thread. We feel that you have an expertise on some subjects but this should not to be spread around all the subjects
I think also that it’s useless to make the mail from Paulus public. It does not help the project. It’s something between you and Paulus and Michael.
Best regards.

Nope as that was it and is it.
Doubt I will be contributing much more and would prefer not to be tagged in future conv.

https://www.kickstarter.com/projects/smartaudio/ankerwork-s600-all-in-one-speakerphone-0

No worries, Don :slightly_smiling_face:
I share the same thoughts about Stuart (not tagging him as he asked not to be). In fact, as I delve deeper into building voice satellite hardware for Nabu Casa, I keep coming across posts from Stuart either here or across Github. It’s becoming clear that I wasn’t ready to hear many of his ideas; I’m still trying to catch up.

That’s not to say I 100% agree with all of the negative things that were said, and I believe it’s that negativity that pushed me and others away. Different people (and organizations) have different priorities and different strengths. Despite being in the voice space for a few years now, I only have a rudimentary understanding of audio engineering. But we’ve made a lot of progress around Rhasspy and other projects despite not having audio expertise, though I do think we could’ve made it easier on ourselves with the right knowledge (but when is that not true? :smile: )


Regarding the topic of this thread, we are discussing satellite hardware over here. While that discussion is focused on an ESP-based satellite, I believe very similar hardware could make a great Pi HAT or USB mic. Specifically, the XMOS XU316 with two microphones, audio out (echo cancelled), and some LEDs.

If anyone knows enough circuit design to build a prototype of such a thing, I’d love to chat with you :slightly_smiling_face:

@synesthesiam

You have to be wary of any hardware or closed software blobs that you can not run data through faster than realtime. As with the size of the datasets that is some process time.
You need to be able to preprocess the datasets of use all the way up the chain with the algs and hardware of use, which is easy enough with opensource software that you just run on a fast computer.

The Speechbrain Sepformer is an example as for it not to work worse than normal they had to finetune Whisper with a Sepformer dataset.

You where involved when Mycroft finally created a XMOS PI hat that was 2 mic. Surely anyone with a Mycroft II and surprised you don’t have one, being a Dev at the time.

I do have a few Mark II’s with the SJ201 daughter board (containing an XMOS XVF3510). It works pretty well for audio, but there are a few things I would want to change. The non-right-angle header and 12V power plug alone will probably put a lot of people off :smile:

Again confused as why would you pick a microcontroller for a maker product that has near zero maker community?
Without research I presume Xmos do provide licensed software that likely is the same that is baked into there XVF3510.
Also from memory as this is how long I have been saying you need to use the KWS to lock the beamform for that command sentence was still missing.
Its still the same as its buying in knowledge because the community lacks the DSP/ML skills to create the essential initial audio processing for a room voice microphone.

The dumbest thing on the SJ201 daughter board was to hardsolder microphones and make placement and isolation a near impossible task.
Like the Esp32-S3-Box the microphones are on a small pcb thats connects by a FPC (Like Pi Cam Ribbons)
Likely 12v is a good input voltage that is a good source for an audio amplifier that is less toy-like, A DC to 5v stepdown is a very common component circuit.

For testing those problems do not matter and maybe some empirical data could be provided.
The AEC on those is pretty good as none linear, but that doesn’t cover the noise by common media such as TV, Music, Radio …

Tha above is the problem as you have that currently with the esp32-s3-box and from the results you are getting due to lack of algs and DSP in the community and anybody who is capable of steering it.

This is where my head explodes in absolute confusion as you do have a Microcontroller that is capable that does have a community, but unfortunately the skills of that community is limited.
There is no problem running a KWS on esp32-s3 but the chosen KWS uses a closed source blob provided by Google and uses layers not supported and likely for any micro-controller it is the same.
The esp32-s3 can very well run KWS and I have said many times that likely a CNN, DS-CNN, BC-RESnet and maybe a CRNN all documented in detail at google-research/kws_streaming/README.md at master · google-research/google-research · GitHub with a training API that also includes tf4micro, as said on many times.

Unfortunately the gimic of custom KW was sold as a key feature as that is something even big data can not afford as KW dictate is due to the datasets they hold.
What dscripka did is exceptionally good to quickly get a KW in operation but sadly no guidelines where given and no option to collect correct KW was in place with an option to forward and send as opensource data.
GitHub - dscripka/openWakeWord: An open-source audio wake word (or phrase) detection framework with a focus on performance and simplicity. is brilliant for that, but is a rather fat KWS for many microcontrollers and actually less accurate than many with dedicated KW datasets as mentioned above.

Swapping to another microcontroller because you lack the tech and steering skills to create a solution is still going to be the same. In fact even worse as the community for support or dev is even more sparse than the current esp32-s3.

The esp32-s3 is an esspressif technology demonstrator where they give a framework, software blobs, working hardware back with a github of PDF circuit diagrams and bill-of-materials and even the PCB Gerber files.
Esspressif does know enough to design a circuit and has and all that info has been available to you for some time…

Likely abandoning optimised C programmed microcontrollers and throwing Victorian engineeering of application SoC’s such as Pi and other will allow the use of hobbyist Python use with much completed permissive opensource ready to be rebranded and claimed as own. Also the algs used can be run on faster computers for dataset preprocessing.

The respeaker 2 mic is still avail and likely and enclosure for a stereo mic ready to plug in would be beneficial to the community.
speed of sound in dry air at 20 °C = 343 m / s
343,000 / 48000 (48Khz common max sample rate) = 7.1458333
So 2 mics centered at 71.458mm likely is a good choice with current ADC’s before aliasing starts to be a problem.
Speex does have AEC and Pi’s have Pixel rings galore… also Openwakeword also fits and runs.

Also a tip the OKdo ROCK ZERO 3W 1GB with Wi-Fi/BLE without GPIO - OKdo is a Cortex A55 that for ML can often outperform the A73 of the Pi4 for £18.60.
Maybe a stereo mic or USB stereo mic as USB2.0 cables can be fairly long and multiples can be used unlike a Hat. Or plug a enclosed mic into an available 2 mic soundcard (plugable or axagon)

Plugable uses http://www.xiryus.com/wp-content/uploads/2021/06/SSS1629A5-SPEC.pdf which I am not sure but could be exactly the same just a silicon relabelling of Cmedia CM6533 https://www.cmedia.com.tw/support/download_center?type=477_datasheet

From Pi to MiniPC can use USB and no driver needed.

Array Microphone -USB-SA Array Microphone by Andrea Electronics have been doing one at fairly silly prices for some time as all it is, is a stereo mic and stereo mic usb sound card. The above 2 chipsets are from the Plugable stereo Mic and Axagon USB sound cards.

I would also be happy to make a PCB compatible with the Pi either via the GPIO header or USB. If I play my cards right, it may be possible to make a device that works a RPI or has a second PCB added to allow the ESP32-S3 to be connected instead.

1 Like

In your opinion what would be an ideal setup for a satellite speaker, be it RPi, ESP32 or a Conference speaker on a 150ft long USB cable. It would be great to develop the next step for the open source assistant community.

I think most consumers want a single point satellite, with only a power lead to it, and we have seen the Google home and Amazon echo’s make an impact on the market, I’m assuming people want something similar to those capabilities, without breaking the bank for the hardware costs (£30 to 100 is a good price range)

Satelite speaker doesn’t really compute with me as what they are is just consumer versions where they have bunged wireless audio, beamform mic and pixel ring into a plastic package.

Wireless audio, beamform mic and pixel ring are all individual elements and there is no need for the term satelite as its purely an enclosure where by choice you might put them together.
A room might have a single Pixel inidicator that is much larger than a consumer unit. It may have dedicated wireless audio as opensource already has 2 excellent wireless zonal audio systems with Squeezelite and Snapcast.
A room for coverage may have several mics to provide total room coverage.

A satellite only exists because Google & Amazon and likes have created them and they like selling multiples of them.
This satellite thing is nonsense to me and the only thing we struggle for in opensource is a zonal wireless microphone system that is the input as a wireless audio system is the output as we already have zonal wireless audio systems (squeezelite & snapcast also they are pretty damn great).

I had a hunch we could have a websockets server connecting all wireless KWS microphones and selecting the highest KW hit sensitivity as the stream to use for that zone.

Consumers have no alternative than single point satelites, because that is what fits the business model of the likes of Google & Amazon.
Opensource should be like HomeAssistant where it doesn’t define and confine what users should have. It creates choice and there is no such thing as a satelite its just that in that scenario you have chosen to sit your mics on top of a very small speaker in a plastic box and if that floats your boat then go with it.

Others may take some Audiophile grade speakers of the latest and greatest in design and wire them up to a wireless audio system that is much more than just a ‘satelite’.
A room might have a bespoke central indicator, or maybe its just a screen that auto-changes channel on use.

Corporates want users to have consumer electronic devices because that is what creates them revenue and they have been dumping for a loss to create a moat for there specific services.
Consumers may want single point satelites, but users of the latest and greatest in home control, audio and visual, may want something like HA that is open and doesn’t define and confine to a single satelite.

We already have zonal wireless audio and pixel rings galore and I think what we are missing are ‘zonal kws’ devices that its choice to how many are in a room and if they are in a single box oddly called a satelite.

Talking electrical with you we just need a decent analogue AGC like on the https://www.aliexpress.com/item/1005006233720383.html
Maybe 2x of them in a ready made enclosure the lineout signal can go for fairly long lengths and is no different to a satelite needing a power cable.

I think those as they are cheap and easily interface to $10 stereo ADC USB soundcards such as the Plugable Plugable USB Audio Adapter – Plugable Technologies

Likely slightly better than the Respeaker 2mic but that is still very valid and all-in-one for $10 especially if it didn’t have a right angle conector and the mics actually pointed at you than the ceiling, but you can always wall mount, than on a table.

The bits we are missing are the algs as Google has had voiceprint voice filter tech for quite some time.
The latest in greatest in binaural source seperation and voiceprint speech enhancement has seen a plethora of papers submitted recently.

I do have my delay-sum hack code GitHub - StuartIanNaylor/2ch_delay_sum: 2 channel delay sum beamformer and was always hoping a proper coder might help Neon optimise the FFT routines.
There are also various filters on github that with a smattering of Rust or C/C++ could likely work on the Pi3 and above sort of level.

Its the software we are missing as its not been implemented and the blinkered focus of satelite with mics ontop of speaker is actually the worst scenario, unless very well engineered like the consumer products we see and that is far beyond simply 3d printing an enclosure as some have done previously.

Beamforming, AEC, Filters and stereo ADC’s from the 2 mic respeaker to various soundcards have always been available as the KWS.
We don’t have any good datasets and that is where the Big Guys have opensource killed.
We should of been collecting them a long time back and for some reason that obvious necessity has been ignored.

There is no ideal setup and we are short of quality building bricks (datasets) to build upon what we have.

True, but a single central device works well for consumers as they take up little space, and combine multiple services. Sticking a Mic on top of a speaker isn’t a logically good idea, but its the most compact way that consumers are happy with.

That could work, though would require the consumer to either have a rats nest of wires coming from a central point, or require multiple Wireless nodes, each streaming a microphone signal back, and then have the end user to map the layout of the Mics to the closest speaker for that zone, all of that adds convolution to the end user who most likeley wants something plug and play

Indeed, there is no limit on how much someone can spend, but there is a reason the Echos and google home have been adopted so broadly, I think is because they are a good enough unit for most users for playing music.

Of course, everything the corporations is done for the most profit/ recurring review, with the minimal work. The amazon echo was intended to sell amazon products and music to the masses, that’s why its sold at a loss and heavily locked down so we cant stream local music. Though as they can see that’s not as profitable as expected, they are dropping services and quality. Most users here have likely come from those ecosystems, due to the reduction of quality of service, and I’m guessing most of them would like a device that can deliver a similar service to what the commercial smart speaker could do in its hay day.

is there a particular good looking project that incorporates both Audio pickup and 5W Audio playback in a single package, that works well for the community projects?

True, but there is only so much the community can do, most of us are doing this for the fun of it in evenings, where as they big corps are doing it with a mountain of cash backing them for potential future revenue streams. I’m not sure if its so much ignored, or if the community didn’t have the skills available, or interest in the non low hanging fruits

As for the hardware, I’m trying to get my head around what you are suggesting. Would you have for example in a living room, a single microphone in each corner of the room, and a central speaker for rendering the feedback? I’m guessing most consumers wouldn’t be happy doing this. Have a look at conference room audio setups, I’ve seen very few of them with remote microphones, most of them have a combined speaker and microphone in a single unit, even on a large desk (8m) they seem to decide that its not worth the mess of wires

Consumers are happy with the highly developed and engineered housings that the likes of Google and Amazon can push out with economies of scale.
The point is when a maker sticks a mic on a speaker in a plastic box, irrespective of AEC that is all it is and nope so far it has made no maker that happy RiP Mycroft.

No we are talking mainly wireless for coverage and this is what the esp32-s3-box is currently doing, streaming 24/7 room audio.
You can on a wireless node have multiple wired microphones, if that helps with coverage, again its choice.
We don’t actually have a system that does far-field well and you don’t have to if you just add more mics to gain coverage.
Some rooms are of a size that a single microphone will never do.
The device I am suggesting is a ‘KWS mic’ that only broadcasts on KW hit and only the device with the highest sensitivity hit is streamed until end of command sentence. How many you have total or in a room is choice.

Exactly its about choice as I have a $10 amp on the back of bookshelf speaker I already had (I have 2 actually as true stereo sounds so much better), or I could put it in a box and call it a Satelite.
The point is as a user I am not defined and confined to only a toy like box for my audio. I don’t have the engineering skills that Google & Amazon have or the economies of sale. So from cost effective 2nd user speakers to the latest in greatest I can have choice, but would unlikely be able to manufacture the likes of a Nest or Echo speaker.

There are quite a few such as GitHub - sonocotta/esp32-audio-dock: Audio docks for ESP32 mini (ESP32, ESP32C3, ESP32S2 and ESP8266 mini modules from Wemos)
I am not a fan and again against bundling all as a single option.
There are many squeezelite boards but prefer the choice of a lineout to give choice over the amplifier I should use and speaker. Also unlike consumer satelites I can upgrade amp or speaker without needing to change satelite.
Also I have preference to Snapcast as it has a much finer sync timings and is multi-channel and could also be your TV’s surround sound. Also it can be any device with choice of any soundcard or DAC for audio out.
Again its purely about choice and not defining and confining to a specific role and letting users decide on what and how they will use it.

I am not suggesting anything, I am saying its choice and users can have whatever setup they wish.
Its you who is defining a single satelite and confining to that definition and keep repeating this strange conception that it can only be so.

I think you might be confusing openWakeWord (which runs on the Pi) with microWakeWord which runs on the ESP32-S3. In fact, microWakeWord is based on the inception model from the Google KWS streaming repo :slightly_smiling_face:

I like this idea, and it could actually be done today with a regular ESP32 (not the S3). Do you think beamforming is less important if you have handful of these KWS mics in a single room?

I saw your code for the delay sum beamformer, and am interested to talk more about it. I wonder if it could run on the ESP32 using the esp-dsp library for FFT. Looks like there’s some potential for a NEON-optimized version using kissfft.

I think this would be awesome. I could imagine even a single device where you plug it into power and it acts as an ESP satellite, but when you plug it into a Pi via USB it acts as a USB mic/speaker combo.

No I am not microWakeWord is new and the repo only a month old and I had switched off way before that.
I don’t know inception or how its being run (streaming or not) it still likely suffers from having very limited datasets (allow users to submit and promote certain KW).
I think you will find they are using keras-applications/keras_applications/inception_v3.py at master · keras-team/keras-applications · GitHub than the GitHub - google-research/google-research: Google Research
A micro KWS could of happened a long time ago if it wasn’t for the lack of datasets. Synthetic datasets work but are a lesser substitute to actual real datasets. If we actually started collating KW with a opt-in we would have so much choice of what KWS we could run.

I think both is likely the best but likely yeah we can bypass technology by simple hard physics of locations and multiple locations.
Using the same KWS and KW hit argmax to pick the best stream.
(All ‘ears’ would connect with a small buffer delay and you just drop connection to the ones not needed).

The delay-sum just hacked portaudio and a circular buffer to robins code and there are many FFT libs that are Neon optimised kissfft or pocketfft could likely speed things up and lighten load even if not that heavy already.
It just uses GCC-PHAT to get the time delay between the mics and then sums the 1st mic with that sample offset.
Its a very simple beamformer but because its a simple sum likely it will have no signature or artefacts that could confuse an ASR such as Whisper.

Maybe but without algs to run more mics means absolutely nothing.
We do have the respeaker 2mic & at least 2 stereo mic USB soundcards.
We just dont have any source seperation algs and just the simple beamformer I did.
Without its just yet another multi-mic recording device with not real advantage.

Use OpenWakeWord and promote 3 KW’s microWakeWord uses Okay Nabu, Hey Jarvis & Alexa
Alexa because its anothers KW is likely a bad idea HomeAssistant, Hey HA … maybe
Add the code to capture the KW and package them and maybe even go back to the days when I was shouting for a Mic word prompter to record datasets.
If done raw on 2 mic or stereo usb even better.
Provide meta data as we don’t need names or address, just native speakers and the region/country they would associate with.
Commonvoice is near useless as it has near no metadata and has more non-native english speakers than English and there is a huge spectral difference.
Age, Gender are also great we just don’t need identity data.

As for creating PCB’s again without the algs/software to run more advanced source seperation and beamforming it prob pointless.

Espressif have already done the circuits of the Esp32-S3-box version.
What would be cool is to hack off all the unessacary of the technology demonstrator as it does have beamforming algs.
Use the 4 channel ADC they use and on the same PCB have another ESP32 running squeezelite where the DAC lineout connects to the AEC ref of the ESP32-S3 ADC 3rd channel.
The rest. screen and all the other bumf get rid so its purely far-field mic with dual micro-kws and attached zonal wireless audio.

I entirely agree with that, the dev kit is just that, for development and POC, its not ideal for the majority of users as a smart speaker

We could use the ESP32-Korvo v1.1 as a base design, and dedicating the ES8311 to one ESP32, and the ES7210 to the ESP32-S3. I’m guessing the Dual ESP32 approach would allow for better handling of services, but would that reduce some of the user experience at initial setup, as I’m guessing both ESP32’s would have to be configured independently (unless we could get them to sync WiFi setup via UART to each other?)

Screen definitely isn’t needed
Battery charge circuit isn’t either
Neo pixel ring would be useful
Hardware Mute would be wanted by a fair few
5-20W mono amp would be nice to make a contained smart speaker

Have a look at the project I’ve been working over on GitHub
it is currently based around the schematics for the ESP32-LyraTD-MSC and the updates of the ESP32-S3-BOX-3, using the ZL38063 DSP at its core. I had been avoiding the ES7210 & ES8311 combo, as I believed that was putting extra workload on the ESP32-S3 that could of been handled externally on the DSP (AEC, NS, BF, AGC functions), which has support in the ESP-ADF. Though the closed source nature and inaccessibility of the tuner software to the public is less than ideal.

I suppose the question I should be asking is (as I’m considering changing the project trajectory), do you think that the ESP32-S3 with a ES7210 ADC has a mature enough software process to not need a dedicated DSP that can handle the AEC, NS, BF, AGC functions?

I don’t know for sure, you are in the UK are you not.
Msg me your address and I will send you the orig and lite version of the esp32-s3-box haven’t got the 3rd version.

I am hoping the BSS (Blind Source Seperation) software is enough. I have hunch what esspressif is doing is using BSS to split sources and then runs KWS on each split stream to select the working stream.
That is the catch with BSS as it will split audio sources very well but what channel they end up in is near random.
Likely if you jetison all that unessacary then there is room for x2 fatter KWS to select the command sentence stream for websockets.

The esp32-s3 with its vector instructions is almost 10x faster than the standard esp32 so also the beamforming code should port over whilst removing portaudio for the esspressif I2S streams.

Also likely we don’t need a DAC on the esp32-s3 and also the squeezelite can be a seperate PCB that is just connected from its dac output to the esp32-s3 ADC 3rd input.

Really all that is needed is the 2 mics, ADC & S3 on a tiny PCB with the 2mics on the same FPC ribbon. I have forgot actually what dimentions and sample rate the audio comes in at on the box design but the details are all there,

Likely it needs some sort of fusion service to commision a small batch run, but what you can do hack a box to do the same with bespoke firmware.

If anyone can hack the Box and get the stereo audio stream offboard as samples its likely from a listen you could tell how effective it could be.
The AEC is likely very sensitive to any latency with a small tail due to its memory constraints and can not remember if that is hardcoded in the blob or can be extended.

Indeed I am UK based, I’m in Truro, Cornwall. That’s a very kind offer, it would be nice to give those devices a look over.

As you mentioned getting audio out of the devices, I would be happy to modify them to provide a 3.5mm audio jack output, though the DAC on both devices will only output mono, but it should still be good enough for testing.
Its interesting to see that the BOX Lite has no AEC feedback from the DAC

I have to admit the software for this project is still currently above my understanding, though I’m hoping that I will be able to learn in my free time, while also trying to make something that is of use to not only my family but hopefully the community too.

I’m coming around to the thought of using a ESP32 for audio playback and a ESP32-S3 for audio capture, it does feel like it adds convolution to the end user for initial setup, but if that allows us to increase the processing power to improve the user experience post setup, so be it. I’m sure if the Proof of concept device works, we can look at reducing costs (possibly by using a S3 for both roles, and following some RPi manufacturing concepts)

Sounds good to me, effectively one PCB would house a ESP32 with a ES8311, a line level (Interconnect) output and possibly an audio amplifier, along with support components
The second PCB would be a ESP32-S3 with a ES7210, with line level input for AEC, a Hardware mute control (using a data buffer) and I would probably aim to populate with 3 mics that the ADC is capable of handling.
Would there be a need for the mics to be on a FPC PCB? I know most smart speakers have the MEMS units mounted pointing at the ceiling, but we could mount them on a sub PCB to get them at right angles pointing outwards around a circle (sorry my ignorance of audio design is likely starting to show)

1 Like

They are in the post.
I suggest you stick to 2mic as with bind source seperation you split into nMic number of streams.
In domestic scenario’s often its just the command voice and a 3rd party source of noise.
Supposedly it does do 3 mic but that means running x3 KWS if I am right about operation.
https://docs.espressif.com/projects/esp-sr/en/latest/esp32s3/audio_front_end/README.html

You really need to use the Esspressif IDF and not the arduino framework but I think it has extentions for VS Code and eclipse that should make a decent IDE.

Likely you will have to hack the AFE where its fed to wakenet
https://docs.espressif.com/projects/esp-sr/en/latest/esp32s3/wake_word_engine/README.html

https://docs.espressif.com/projects/esp-dsp/en/latest/esp32/esp-dsp-benchmarks.html

Likely it will as the optimisation for the S3 is already done so its just a matter of porting the FFT code.
FFTs 16 bit Fixed Point which the audio is, is much faster on a S3
The S3 has 2x I2S ports that are bi-directional so you could use a quad ADC with 3 mics and a broadside array will have more side attenuation and still have the ref input.
The only real load is the GCC_PHAT where it tries to correlate the time position of the samples in 2 channels.
If you do that once then the rest of the calcs should be much simpler trigonometry and maybe you could create a x4 mic beamformer running Gcc_Phat on x2 mics only as you know the geometry of the x4 mics and the direction of sound.
A x4 mic square is essential 2x broadside arrays that could make a final endfire with only a single Gcc_Phat calc of x2 channels.