Just saw this:
Not much time to read it all but I guess it will interest lot of people here 
Just saw this:
Not much time to read it all but I guess it will interest lot of people here 
Oooooooooooooooooook
One more wake word engine to forget:
All models generated with Picovoice Console expire after 30 days. To generate models with longer expiration dates, a distribution license is required.

Porcupine is the datum to head for and dev target to aim at. Its considerably lighter whilst being more accurate than anything else.
The porcupine offering of several predefined keyword offerings or a custom training with a distributed license is required as no-one is going to want to retrain every 30 days.
Having said that custom wake word need is a gimic as do you call a dog a cat or Paris, Quarkybeddingspot because Paris is that Biatch who broke your heart?
Nope a dog is a dog and Paris is Paris and both Google and Amazon have been extremely successful running without custom wake words.
I am sure there will be many having a tantrum that they wish to call their AI Pretty Petal or whatever but when you have clear defined projects such as Rhasspy its pretty obvious what the KW should contain, but often like the commercial guys a couple of alternative models are offered.
Google & Amazon make use of that limited choice because it enforces use in numbers and with their analytics as those numbers feedback and improve the model.
For small projects custom keywords can dilute numbers so much that data availability becomes too sparse for effective use.
Custom keywords are not necessary, can actually hinder a project, but hey lets pander to those who want to give the puppy a name.
I have heard what @synesthesiam is doing with TTS which looks like its something Picovoice do which makes great models for âNews Readerâ middle of the road perceived English on good TTS engines.
The simplest way with the smallest model achievable and very high accuracy is to train with your own voice.
A broadening of size would be to train with a dialect that you use which will give much higher accuracy than general models.
If its single gender use then gender voice training also makes big increases in accuracy.
The required wealth of data is likely never going to be harvested when a small project is diluted over many key words.
For the âPretty Petalsâ out there custom keywords maybe seen as a value feature and some others might not care and have more qualms about accuracy.
What is interesting with Picovoice is their validation dataset as if they used TTS then their accuracy results donât count for jack with real voice especially with strong dialect representation.
It grates on me that regional dialect and gender difference are excluded from datasets whilst a middle of the road elite are happy that it works great for them.
That dialect, gender even age and especially âown voiceâ are excluded from generalised datasets or datasets are available generalised where local harvesting will create periodic model retraining.
Even opt ins to provide not just the data to opensource source datasets but the meta-data that is crucial for tailored training accuracy.
If picovoice is offered with a choice of non expiring keywords then it really doesnât matter about custom keywords unless your trying to sell for commercial use, which they then provide a commercial license for.
Custom training rather than blackbox models is very important and if we are going to be inclusive and provide dialect, gender and age specific or weighted datasets so all can enjoy maximum accuracy and small lightweight models.
Picovoice isnât opensource its shipped as binaries and that is the only reason why for many its one to forget.
I for one will be really happy when some realize the voice .AI goldrush bubble has already burst and focus is purely on some really good opensource.
This was already the case, but accounts only for custom wakewords
Ah yes I was mixing with precise lol !!
Anyway custom wakeword are everything. Google and Alexa doesnât offer it only for marketing reasons. Having millions people saying ok google make it a standard and sell it to everyone. With custom wakeword no one could know which assistant you use. Purely commercial.
Really a shame snowboy is ending, I have better results with it than with snips wakeword which have ever very good. Custom with each for comparison.
There is a new opensource KWS as I also find Precise a bit heavy and not that precise.
Has model generation utils and if @synesthesiam does his TTS KW dataset populator the model generation tools are already there.
Also utils to capture own word and bolster model creation.
https://github.com/linto-ai/linto-desktoptools-voxharvest
I think it was all a rapid dev project approx 6 months ago and still has some rough edges but looks really promising.
Who and how someone would know what assistant you use in comparison to accuracy is totally inconsequential as who and what would know?
It wasnât just for commercial reason they where picked for uniqueness of voice capture and sylable count and there data capture fed back into their models.
If Alexa can respond to âcomputerâ where is the commercial identity you mention.
Strangely Google have the non commercial âHey Boo Booâ where ever that came from!? âYabba dabba doâ!
I am waiting for @synesthesiam great new TTS dataset populator and going to augment with âown voiceâ and dialect/gender extracted ASR dataset words and run through Linto-HMG as been playing with it and the results seem extremely good.
But waiting for some rough edges to also be smoothed but herd datasets from small communities are a much better way to go as we can share common opensource voice data for accuracy.
But if you wish to go the route of custom KW then you can.
Actually wakeword is what I miss for rhasspy. Still use Snips on my prod due to this. I though snowboy was the answer, but investing on EOL solution is not something I like.
I need pi3/0 wakeword service for three different custom wakewords (one per person here). Snips is working nice, snowboy near perfect, now whatâs next ⌠
I havenât really tested the https://github.com/linto-ai/linto-command-module or its Lib https://github.com/linto-ai/pyrtstools but it does look good and could fill that gap.
The only model available is French and me being typically monolingual means I am waiting to build a model.
You can try out with HMG now but all the tools are not really available but sound like they could arrive pretty soon.
I have given HMG a whirl with the Goggle command dataset using âvisualâ as a KW and it seems quite promising and also an eye opener to supposedly validated datasets as there are some poor entries in the google dataset that have effect.
Many of us have different perspectives and needs and I will be looking at very simple linto based sateliites probaby feeding a Rhasspy or Voice2Json server.
Its been the KWS that has been my hold up and the models for the KWS but hopefully that is soon to change.
Tensorflow-lite will run on the zero https://www.tensorflow.org/lite/guide/build_rpi but to be honest I think the Pi3A+ is a better option as then I can also run AEC (Echo cancellation) of playing media and allow barge in.
Its ÂŁ10 more than a Pi0WH but the performance jump is approx 10x+ that of a zero so 10x for ÂŁ10 is pretty damn cheap.
Iâve tested Picovoice Porcupine and the results are not that good. 
The CPU usage is awesome and the accuracy is as good as advertised 
BUTâŚ
The search for an easy to create and use wakeword continuesâŚ
I still use Snowboy, works best for me even with custom wakeword.
well i use snowboy too since the other hotwords are not as cool as âjarvisâ
iâd like to have a âvikkyâ hotword ( like the AI in the âi, Robotâ movie)
I know but in 6 months this will be shutdown.
What will happen then?
Custom wakeword for children will expire as their voice changeâŚ
What is I want to add a new wake word later?
I do not think Snowboy is a viable solution as of today (maybe they will release their training stuff⌠one can hopeâŚ).
I do not know 
But the custom wakeword I created (47 yrs old), also work for my kids (6 and 9)
It might not be a good solution in the long terms, but it still works today and for me better than porcupine.
I thought the Linto KWS would be right up your street it being French based is why I am waiting for a dataset populator with my Brit twang.
Not sure if the KW is just Linto on the only tflite model they supply and do you say âHeyâ or â âCoucouâ.
Its been a strange couple of months since my curiosity perked with Mycroft as these projects have been going for such an extended time that it seems odd there are such obvious holes missing from the project.
Seems here if it has Hermes in the title then it will become a community project if not then its something for @synesthesiam to provide.
Irrespective of any KWS we are lite on dataset creation and collection tools and quite a few easy solutions exists but seem to garner little interest prob due to lacking the Hermes tag.
There is a huge wealth of data out there where word extraction from ASR data can provide extremely large KWS datasets that its possible to cherry pick the metadata and create extremely accurate weighted language datasets with at least region and gender. Age does exist to a much lesser extent.
That a huge amount of effort can be given to the completion of skills whilst the very fundamentals of KWS have obvious need is extremely indicative and a sad reflection of âcommunityâ priority vs self.
I have been posting for a while now examples that tensorflow KWS using common tools such as Keras is actually quite easy even for the like of me.
Been posting consistently that the problem is a not really the lack of a KWS its the lack of KWS datasets.
Once more it seems to be left to @synesthesiam to hand out code like Jesus does bread.
I donât code as a pretty good ex hacker in fact not even that level, more molester of code I was pretty reasonable but MS brain damage just makes it too frustrating to learn again.
There are some extremely competent coders here and I am sat bemused looking at some obvious needs and a lack of any input.
I have actually found that Linto more or less have an opensource package for almost every need I can think of, its only @synesthesiam and the TTS2Dataset that they miss.
Its all rather fresh and some rough around the edges but a large wealth of code is already available.
For simple voice data collection they have https://github.com/linto-ai/linto-desktoptools-voxharvest that prob needs some custom fields for meta-data where things like region, gender and age can be added.
Prob needs 2 modes of operation as word based KWS and sentence based ASR prefer slightly different datasets. Both need to be split into folders of along the lines of meta-data and KWS its also handy to create folders for that word.
Own voice can greatly increase accuracy of KWS or ASR and we even lack those simple tools.
That voice workload can also be greatly reduced by pitch shifting and noise addition where a single word recording can become many.
The above is just one example that has pressing need where a herd shared piece of opensource would be extremely beneficial to a number of projects without merely appropriating code, adding a smattering of custom code and branding as your own.
For the preference of word based KWS there are absolutely massive datasets available where it would be extremely beneficial to strip words from ASR datasets and organise and collate metadata.
Its also extremely likely KW can be created by concatenation of extracted words and massive volumes of non-KW could be collated.
Then once more pitch-shift and noise-addition can greatly multiple individual KW/non-KW collection.
I have already posted that Linto have a great KW model creation tool https://github.com/linto-ai/linto-desktoptools-hmg and once again prob could do with a few additions as wish once you have created a folder based dataset you could export that dataset json, for example.
I couldnât give a damn if its Linto or not as I was searching and realizing much was needed and then found Linto already had a headstart. I am having to reuse as its not just creation its maintenance so existing projects have even more value.
Is it so painful that a repo doesnât contain the word Hermes or that another Voice project might have much to share?
Why the search continues, why some solutions are not adopted now, why not collaborated on now for some of us is not such a mystery, but it is extremely frustrating and it negates generally from the project.
If you have @synesthesiam tts2dataset populator then that fills a last gap and once more I have tried to highlight the rest of the need already has solutions and just needs a little additional collaborative work.
By crikey its Hermus!
Your tts2dataset utility should be great but it limits itself by the technology used and creates a narrow dataset on that technology.
We should be able to share datasets and all we need is a central database that can store URLs of the datasets we supply.
It should have queryable metadata of Language, Region, Gender and Age that archive URLs can be submitted to.
I can get 15gb free on my Google drive and could go to work on stripping the ASR sentences of the likes of https://openslr.org/83/ and CommonVoice via the metadata into words.
I could create another account and sign up to a different service⌠Add more.
Each unique subset I create is an archive of root language folder with corresponding metadata subfolders, containing the word folders.
Some fields about source dataset readme.info ting.
An extremely small database and web app could link huge numbers of community supplied dataset archives that can be quickly returned by a metadata query of the key fields you select.
I can just submit my archives that just become part of large distributed dataset.
A login and feedback rating system should also be able to create a sort order.
Output could be even a simple cli text file of wget urls and that is it.
Then we all can start to submit KW datasets or do we just pray to Hermus?
Someone might even grab Julius and convert into Phenomes https://github.com/julius-speech/segmentation-kit same metadata and folder structure containing Phonemes rather than words and maybe that might even prompt an extra query field.
Who knows but its actually that easy.
I guess even
Would be cool if we could go to porcupine 1.9. They added a bunch of new wakewords: https://github.com/Picovoice/porcupine/tree/v1.9/resources/keyword_files/raspberry-pi
However those are not compatible with 1.8.
Porcupine is a great lightweight KWS sadly not opensource but really good and those new wakewords offer more scope and guess for a wake-word they can be more multi-cultural than ASR needs.
Its a small selection but quite a good selection of wake-words. " Added Alexa , Computer , Hey Google , Hey Siri , Jarvis , and Okay Google" all well known wake words
Thanks for the update as that is interesting as it might not be custom or opensource but is relatively accurate.
porcupine 1.9 would be a good upgrade as one of its downsides was the limit on models avail.
Does pip install pvporcupine upgrade it as has the api stayed the same?
Anyone running 1.9 as seems quite a lot of run time optimisation / accuracy improvements have been added in the last couple of releases and interested to findings as it always was light & accurate.
I guess they keep updating the neural nets as they also keep getting optimised and smaller, sometimes hard to keep up.
The main problem I have with Porcupine is that it only handles english phonemes. Pronuncing âAlexaâ or âJarvisâ with a french accent does not work very well.
Oh and it is not âopen sourceâ either⌠
I can confirm, that the new models donât work with the porcupine version included in Rhasspy 2.5.5.
Did you manage to get it working?
Or is an upgrade already planned in the next Versions? 
I agree with @fastjack , Porcupine does not work well with french accent.
So I use mycroft-precise train with my voice on the word âmaitre yodaâ but it has too many false positive. And it detects only my voice due to a lack of other voices in the training.
Dunno havenât tried 2.5.9 update as thought it was supposed to be.
You really need a lot of samples so its like a label in the Google command set (2k+).
If you can record 50 samples you can quickly multiple that by using some pitch and padding changes with a tool like sox.
Then add noise to your samples but would have to check how precise deals with noise as that may cause a double dose.
But yeah with only your voice it will detect on voices which are so similar that it might only be yours (do you do impressions
).
You would be better just picking any dataset such as https://drive.google.com/open?id=1-kWxcVYr1K9ube4MBKavGFO1CFSDAWVG use the google command set for an equal amount of non keywords then add a lot of your voice so its weighted to you but not uniquely you.
Its real problem as even with English word datasets are much rarer than ASR sentence ones.
I am not a fan of Precise so donât use it but the above is a rough guide how to train any word model.
If you can use all your words as their label with a simple dummy model and run through and have the model delete the really low scored then retrain with precise.
Nijmegen Corpus of Casual French: 35 hours of high-quality recordings featuring 46 French speakers conversing among friends, orthographically annotated by professional transcribers.
French Single Speaker Speech Dataset: CSS10 is a collection of single speaker speech datasets for 10 languages. Each of them consists of audio files recorded by a single volunteer and their aligned text sourced from LibriVox.
Traitement de Corpus Oraux en Français (TCOF): Over 500 transcriptions of 124 hours of spoken French. The corpus is divided into two main categories: adult-child interactions (children up to 7 years old) and records of interactions between adults.
VoxForge: Set up to collect transcribed speech for use in Open Source Speech Recognition Engines, VoxForge contains 37.5 hours of oral recordings of texts in French.
??
Ask another French speaker if they will record 20-50 of your KW as again you can multiply with Sox but donât weight the model too much to them.
I am having a go with https://github.com/prosodylab/Prosodylab-Aligner to extract words from as many ASR datasets as possible into word folders and hopefully collecting a qty of each word. (Yoda may be a problem
) I bet if you scoured some vids you could create a collection then its just ( Master? apols my French )
https://github.com/prosodylab/Prosodylab-Aligner is OK but think I need to combine it with a good VAD which maybe I can us the nemo VAD for https://docs.nvidia.com/deeplearning/nemo/user-guide/docs/en/v0.11.0/voice_activity_detection/tutorial.html
Havenât got to implementing the VAD as the aligner does work but still needs work where to cut and maybe the VAD may help.
Definitely do duplicate your set with added noise as otherwise you will get a model that will only work in a quiet room but not when there is any background noise. Duplicating with added noise will give you a much more robust model. There is a tool included in precise to do this but itâs not great (call precise-add-noise --help from the venv for how to use it).
I personally use a simple bash script:
#!/bin/bash
NOISEDIR=$1
DIR="$( cd "$( dirname "${BASH_SOURCE[0]}" )" >/dev/null 2>&1 && pwd )"
for f in *.wav
do
NOISEFILE=$(find ${NOISEDIR} -type f | shuf -n 1)
sox -m $f ${NOISEFILE} noise.$f trim 0 `soxi -D $f`
done
Prepare a folder with lots of pieces of random noise (they should be longer than the wakewords). Save the bash script as addnoise.sh to your folder that has you wake words. Run the script with the path to the noise folder as an argument. This will create a copy of every wake word sample with added random noise from the noise folder.
I will let you describe the SNR levels you should use as you use precise and sure you do.
I actually use them one to one as for a lot of the noise in my case both the wake word and the noise sample are recorded on the same microphone and i get a realistic representation how a noisy wake word will sound in that room.
That is not very representative of the levels many KWS can cope with noise usually the best have a max SNR of 5db.
So at best normalize your you noise 5db lower than you KW.
I just noticed you where not normalizing in your script so you could have a noise file louder than the KW and that is just plain bad.
The predominant signal should be your KW especially if your adding an equal duplication again or your accuracy will obviously plummet.
I am not a Mycroft fan and donât go near but a good KW system it would be automated and also it would have steps maybe 5db & 15db below that to create different levels of noise and save you the hassle and headache.
If you can record the noise you have and pick noise files that might be common noise but hey noise is random.
Also just for others do not add the noise files that you use as noise to your ânot keywordâ samples.
Well it worked for me so far but i do go through the created noisy samples and pick out the ones that just plain dont make sense snr wise.
I am not sure what that means but you just mixed 2 signals one is signal and one is noise hence signal to noise ratio.
So when you mix you should first normalise all your KW samples -1 to -3 is usual and then your noise samples -5db below those or more and maybe do half and the other half @ -15db of whatever you set your KW samples at.
Im just saying it worked for me just plain mixing them. If you normalize all samples than probably the input audio when using them would have to be normalized too. I actually use samples recorded at several distances with several microphones which are not normalized at all.
The input is normalized and yes it matters as with models garbage in is an old adage.
I just had a look and the noise function for precise looks pretty robust I suggest you use than the manner you have just described.
The problem with the precise add noise is that it doesnât give a good spread of noise from your noise data when you have a lot of it. Id rather add normalizing to the bash script.
Feel free to adapt and post the script to how you think it should be.
So you would recommend to also normalize the clean sample set beforehand?
So for example normalize wake words to -10dB and added noise to -15?
You have to normalise to a set value so you have a datum to work from so you know what you are mixing at.
Otherwise you just have not got a clue what ratio they are.
What do you mean it doesnât give a good spread of noise from your noise data?
I havenât used it but also on train is it not called like I thought it might or do you have to call it?
:-if --inflation-factor int 1
The number of noisy samples generated per single source sample
:-nl --noise-ratio-low float 0.0
Minimum random ratio of noise to sample. 1.0 is all noise, no sample sound
:-nh --noise-ratio-high float 0.4
Maximum random ratio of noise to sample. 1.0 is all noise, no sample sound
So at guess at a glance at the code its dependent on how many noise samples you have and the inflation factor? Dunno I havenât used it.
But seriously what you have been doing mixing 50:50 on blind db level files is just the worst method of any as its pure chance what your SNR of KW to noise is.
Its all a catch-22 but just adding more noise at higher levels isnât going to make things better and likely worse.
What happens is your KW image in the model becomes so blurred that the cross entropy with non keywords is just likely to raise. That will produce a lower confidence level and you could be doing the the opposite of what you are trying to achieve.
Your KW always needs to be predominant and clean should be the largest amount of a single block of samples.
You should then split your noise in say 15/15/15 5db, 10db, 15db noise and 55% clean or somewhere around.
You can do it your way but I have noticed it only takes a small number of aberrations in you KW files to actually have a large effect on accuracy.
It will only use very little of the data so the noise added will not be very diverse as its only from a very limited number of the noise files.
Thats why i wrote the script above which will choose the noise file to use randomly for each wake word file from the several thousand noise files i have.
What do you mean by this?
You can simply add a -v (for example 0.5) argument before the noise file in the sox command in the script to adjust the ratio of the noise file relative to the wake word file:
#!/bin/bash
NOISEDIR=$1
DIR="$( cd "$( dirname "${BASH_SOURCE[0]}" )" >/dev/null 2>&1 && pwd )"
for f in *.wav
do
NOISEFILE=$(find ${NOISEDIR} -type f | shuf -n 1)
sox -m $f -v 0.5 ${NOISEFILE} noise.$f trim 0 `soxi -D $f`
done
Yes but as I said when only using a clean set to train with precise you will get a model which doesnât react at all in noisy environments.
Look go and start reading about cutting edge KWS I am not going to use Precise but have have played with quite a few models and what I am talking about is what I have seen from the like of Google & Nvidia.
I have no idea we have such a collection of crap KWS and methods when cutting edge is opensource and documented but all I suggest is to follow their methods not mine, mycroft or the dross we have.
There is just about every cutting edge network in here with results and example code and methods.
Check what they are doing and also in the scientific papers published.
They give you a headache but are extremely comprehensive but you seem to have some false assumptions about models and how things can work.
I agree im not a general expert on key word systems like you in any way. I dont even claim do be intelligent enough to understand any of the machine learning parts happening in more than the broadest strokes.
My advise is simply what i found worked best for me to build models that are robust enough for daily use with precise and no other system. This advise comes from training over a 100 iterations on a few different models and trying what worked best, fine tuning the precise training settings and seeing what didnât work at all. Experimenting on what improved the outcome. And adjusting accordingly. Thats all. So its purely empirical and very focused just on precise and in that only on what gave me the most versatile model in daily use through trial and error.
With my method described i got a model which we now use 24/7 which in our household gives about 1 false positive an hour but at the same time allows me to get a response even when the tv is running with both the 2mic and an electret per your method.
So thats all i claim is my knowledge.
As i donât understand what im doing and at the end of the day i am just another dumbass maker take my experience with a grain of salt and im always happy to learn how i can improve my models further.
I am no expert but you wouldnât mix a cocktail with the blind qtyâs you displayed on mixing noise to your KW model.
as i said i actually listen to the outcome and sort out files where the keyword is apparently over powered. I will experiment with the -v volume factor to ensure that the wake word is always at a higher ratio. Maybe im just lucky with my set when it comes to precise.
Have a play with this and give a go its a simple CNN so its fast to train let it drop out with a patience of 5 or 10.
Test your hypothesis but from what I have seen what your doing looks flawed but hey.
Apols about the code but been aware the Mycroft Sonopy lib is likely broken so been interested in other libs and now all the frameworks are including MFCC math into the framework.
Was just to have a look see and if MFCC made much of a difference to spectrogram but try out some noise and do a cross entropy check on the labels and see how you go.
There is a cpu version of tensorflow 2.4 here
I will have a look at it. I will have to find some time to install everything though.
I dont have a windows or linux x86 here.
Ill have to install python and tensorflow on my mac.
Do you know if it would run on pine rockpro64 or a raspberry pi 4? Because those are the development machines i have set up.
Donât do training on a SBC full stop unless you have days to spare.
You may have to compile as with 2.4 on a I5 3570 the default wheels fails as I donât have cutting edge AVX-512.
Also the python and GCC compiler can often be different.
The compile is actually really easy and for me the only confusing part of the tensorflow info was.
âbazel build [âconfig=option] //tensorflow/tools/pip_package:build_pip_packageâ as "bazel build "âconfig=opt //tensorflow/tools/pip_package:build_pip_package was all that seemed to be needed for cpu base.
The compile is painful though like several hours set it up as a job whilst you sleep as if it fails at least its not stolen your computer for that time.
gpu is a doddle just the same have the cuda and cudnn stuff preinstalled and its âbazel build --config=cuda --config=opt //tensorflow/tools/pip_package:build_pip_packageâ instead
Bazel is relatively easy to install or you can just download the lastest balisk and create a symlink of bazel
Pytorch has annoyed me though as the strong input for torchaudio enforces non gpu mode of intel_mkl math libs or gpu math libs of nvidia and currently on arm the standard libs of openblas & 3wfft and likes have been fired off into the either for hardware specific libs!
I would really be having a good look at the Nvidia Nemo framework but should of known better with Nvidia even if its supposedly stamped as âopensourceâ
The scripts here are pyhtonically atrocious but it was just me having a look at some of the additions of tensorflow 2.4
I noticed that tf.signal now has internal audio math so either use the collab or my python adaptation.
Which is this
Also they have added mfcc to that math
So I just hacked that in
Really fast simple and ultimately useless models but extremely good to assess things due to build speed and if you can place all dataset into labels of there own id with sufficient samples as then you can check how much cross entropy they have.
Where running through you dataset samples on the model you train and deleting the low dross start < 0.1 delete retrain. Delete < 0.4 or 0.3 retrain and you will find an accuracy increase of 3-4%
sample_file = data_dir/'go/3d53244b_nohash_1.wav'
sample_ds = preprocess_dataset([str(sample_file)])
for spectrogram, label in sample_ds.batch(1):
prediction = model(spectrogram)
print(f'Predictions for "{commands[label[0]]}"')
print(commands, tf.nn.softmax(prediction[0]))
Loads up a single train wav and runs inference on the model and the softmax for all labels is shown and hence cross entropy is checked.
Your just using a simple an quick to train model to test a dataset before submitted to the format and train of a desired model.
I often use âgoâ as if things are going wrong its often a quick canary due to its simularity to ânoâ and it shows.
MFCC just complex the non timeline axis of the model to 13 and greatly reduce the resultant model parameters.
The image is resolution is timeline x 13 so the remainder is focussed on and gives an accuracy boost over spectrogram of a couple of % but its biggest advantage it does this whilst compressing the model.
The high frequencies are just thrown away as to the ear they are extremely low energy and no use in general for recognition.
The reality is though we should be expecting anyone to do any of this S*** or even have a care in the world as from datasets checks to noise addition this should be all part of automated tools which we have a complete lack of any.
Once more the Linto HMG is the best click and view model generator I know but from datasets to tools to KWS rhasspy either omits or what is available is lack lustre and its not because of lack of mention?
@JGKK PS out of interest I thought I would try the full tensorflow on the Pi4-2gb I have and yeah its 300% slower than my I5-3570 but actually its a lot faster than I thought.
https://github.com/bitsy-ai/tensorflow-arm-bin as the downloads on the TensorFlow site are old and also I think they have posted the armv6l one twice
If you just want to fire off a train and forget actually yeah you could train on a Pi 4.
Test set accuracy: 89%
Predictions for "no"
['up' 'down' 'stop' 'no' 'go' 'left' 'right' 'yes'] tf.Tensor(
[1.4503686e-04 4.8033642e-03 2.9509642e-05 9.7052008e-01 2.4486762e-02
2.4404756e-06 6.0085756e-07 1.2100643e-05], shape=(8,), dtype=float32)
Predictions for "right"
['up' 'down' 'stop' 'no' 'go' 'left' 'right' 'yes'] tf.Tensor(
[1.1338150e-14 1.5254406e-15 9.5040550e-19 1.6817958e-17 5.2986136e-17
9.5589510e-09 1.0000000e+00 2.3156420e-16], shape=(8,), dtype=float32)
Predictions for "left"
['up' 'down' 'stop' 'no' 'go' 'left' 'right' 'yes'] tf.Tensor(
[9.3506897e-06 2.2647366e-06 2.0204541e-05 4.9854199e-05 2.3224919e-07
9.9054682e-01 1.4057276e-05 9.3572428e-03], shape=(8,), dtype=float32)
Predictions for "go"
['up' 'down' 'stop' 'no' 'go' 'left' 'right' 'yes'] tf.Tensor(
[6.7035542e-09 1.8232617e-04 9.5976709e-04 9.9639314e-05 9.9875832e-01
2.8040101e-12 6.1784431e-09 5.8221472e-10], shape=(8,), dtype=float32)
Run time 841.416019846045
14 mins default pi4-2gb no OC
python3 simple_audio_mfcc_frame_length1024_frame_step512.py