Data processing agreement
What happens to your players' audio once it leaves your machine, written against how VoiceSniffer actually moves data rather than against a template. Applies to cloud mode only.
What changed on 5 August 2026. Section 4 now describes the local debug capture setting, which the previous version's claim that there is nothing to turn on was wrong about. Section 6 is new and covers the Discord webhook. Section 7 replaces an opt-out promise the plugin does not keep. Section 12 is new and covers automated mutes. The sub-processor list moved to Annexe B and now covers everything, not only the audio path. Nothing about audio retention changed, because that position has never changed.
This agreement covers what happens when Sniffer Studio processes voice data on your behalf. It applies to the cloud mode of VoiceSniffer only. In built-in and docker modes we never receive your players' audio, so there is nothing for us to process and this agreement does not apply to you.
1. Who is who
You, the server operator, are the controller. You decide that voice chat on your server is moderated, you choose what happens when somebody is flagged, and you answer to your players for it. Sniffer Studio is the processor: we transcribe and classify audio on your instruction and do nothing else with it.
If your players are in the EU or the UK, you need this agreement in place before you switch cloud mode on. That is Article 28 of the GDPR, and it is your obligation as controller rather than a formality we are adding.
Your documented instruction is the acceptance recorded in the panel by the signed-in account that owns the licence, together with the settings you configure. No server can be paired to cloud mode until that acceptance exists, and a paired server stops renewing its token if the account that owns the licence has not accepted the current version. That is enforced in code, not by policy.
2. What we receive, and only that
In cloud mode the plugin sends us one short Opus-encoded audio clip per utterance, together with:
- a random request identifier;
- the player's Minecraft UUID;
- a language tag, or
auto; - your server's bearer token, which identifies your subscription.
We do not receive player names, chat messages, IP addresses, positions, or anything else about your server. We never ask for a player's real name, email or age, and you should not send them. The request schema is closed: a field we did not ask for is a rejected request, not a new column.
3. What we do with it
Each clip is decoded, transcribed by speech recognition running on our own machines, and matched against moderation rules. We return the transcript, the rules that matched and a severity. Then the clip is gone.
Audio is never written to disk on our side. It exists in memory
for the duration of one request, typically well under a second, and is not logged,
cached, queued to storage, or used to train anything. The container the processor
runs in has a read-only root filesystem, so there is no path it
could write a clip to even if it tried; the only writable location it has is a
64 MB tmpfs, which is memory rather than a disk, and nothing in
the request path uses it.
Transcripts are not written to our logs either. When a request fails, what goes into the log is the request identifier, the server, the language and the exception. The transcript and the request body are deliberately not variables on that stack, so a stack trace cannot carry them.
One qualification, because "then the clip is gone" is about audio and text behaves slightly differently. The verdict we return, which contains the transcript, is held in an in-memory duplicate-request cache for 120 seconds, so that a plugin retrying the same utterance gets the same answer instead of being charged for a second transcription. That cache is memory, it is scoped to one customer and one server, and 120 seconds is the longest any text lives inside the processor itself. Nothing writes it anywhere.
We do not send your audio to any third-party speech or AI service. The recognition model runs locally on hardware we control. That is a deliberate design choice, and it is the main reason this agreement is short.
4. Audio on your side, including one setting that records people
The paragraphs above are about us. On your own hardware, one setting changes the answer, and you should know it exists before somebody finds it in your config file and asks you about it.
capture.debug-wav-dump in advanced.yml is
off by default. Switched on, it writes every captured
utterance to your server's disk, as a .wav and a matching
encoded file, under debug/utterances. It exists so that a broken
install can be diagnosed. Nothing deletes those files, nothing caps how many there
are, and they are recordings of your players.
Each file is named with the time and the player's Minecraft UUID, so the directory is not an anonymous pile: it is a per-player archive of recordings.
If you switch it on: that is you recording people, on your hardware, as controller.
It is not covered by this agreement, because the files never come to us, and it is
not covered by your players' expectation either, because the join message the plugin
ships says audio is not stored and does not change when you turn this on. Change the
message or do not leave the setting on. Turn it off and delete the directory when
you are finished diagnosing. The plugin does warn you: it logs a loud line at
startup while the setting is on, and /vs status reports it.
Your server log, separately. Whether or not the dump is on, every flag is written to your console and server log at INFO, with the player's name and the transcript. That log is yours, has no retention deadline, rotates into archives that persist until deleted, and is the file operators most often paste into public support threads. It is not something we receive, and it is worth a moment of your attention before you paste one.
5. What we keep, and for how long
5.1 Counters, per server per day. How many requests, what they came back as, how many landed on each severity, and how many seconds of audio that adds up to. Numbers only. There is no field in them that could hold a word somebody said, and the schema that receives them rejects one that tries. Kept indefinitely, because they contain no personal data and they are how "how much has this server used" has an answer.
5.2 A moderation record, for flagged utterances only. When an utterance matches one of your rules, we store:
| Field | What it is |
|---|---|
| Player UUID | The Minecraft UUID, which is all we are ever sent. Not a name. |
| Server | Which of your servers it happened on. |
| Timestamp | When it was moderated. |
| Rule, category, severity | Which rule fired and how serious it is. |
| Rule count | How many rules that one utterance fired. |
| Request id | Makes the row addressable later without carrying any of its text. |
| Matched text | The part of the utterance the rule matched, truncated at 200 characters. |
| Transcript | What was said in that one utterance, truncated at 1000 characters. |
Nothing at all is stored for speech that matched no rule. That is almost everything anybody says on your server, and it never becomes a record anywhere. This is a moderation log, not a recording of a voice channel.
Thirty days, then deleted. Every record carries its own deletion deadline, set when it is written rather than computed when it is read, so changing the window cannot silently extend records already stored. A clock ahead of ours cannot buy extra days: a timestamp in the future is anchored to now. A sweep inside the service deletes expired records on a timer, and a record past its deadline is already unreadable before the sweep reaches it. Thirty days is not a target, an intention or a maximum. It is meant to be enforced by the thing that stores the record.
Why we keep it. So that you, the controller, can review what was flagged on your own servers when a player disputes a mute: who, when, which rule, and what was actually said. Our lawful basis for processing it is your instruction as controller. It is visible only to the account that owns the licence the server is paired to, and that scoping is enforced twice, in the database query and again before anything is serialised. We do not read it, we do not analyse it across customers, and we do not use it to train anything.
Still never audio. The record above is text. No clip is kept for any period, for any reason.
6. The Discord webhook, which is yours and not ours
VoiceSniffer has an optional webhook.url setting, blank by default. If
you fill it in, then every time a player is flagged, an embed containing the
player's in-game name, the severity and what was said is posted to Discord,
by your server, directly. Audio is never sent, but the transcript is.
Note the name. Everywhere else in this system a player is a Minecraft UUID and nothing else: it is all the plugin sends us, and it is all we store. The webhook is the one path that carries a player's actual name alongside their words, and it is a path you switch on rather than one we run.
That transfer is yours, not ours. It does not pass through us, we cannot see it, and we cannot control it. What it means is that by setting one line in a config file you have made Discord a recipient of your players' speech, and you have made everyone who can read that channel a reader of it. Point it at a private staff channel, keep the membership of that channel tight, and remember it when you answer a subject access request, because those messages are part of your answer and they do not expire.
7. Security
Specific commitments rather than adjectives.
- Transport is HTTPS. The plugin refuses to send audio over plain HTTP to anything other than your own machine or your own private network, so cloud mode cannot be misconfigured into sending audio unencrypted.
- Your bearer token is the only credential. It authorises your server and nothing else, it is compared in constant time, and it is never written to a log.
- The processor runs as an unprivileged user, in a container with a
read-only root filesystem, all Linux capabilities
dropped, no new privileges, and a process limit. Its
only writable location is a small in-memory
tmpfs. - The processor's own port is bound to loopback only and is not configurable to anything else. It is reachable from outside only through the tunnel described in Annexe B, so there is no inbound firewall hole to get wrong.
- Request size, clip length and frame count are capped, and memory and CPU are bounded, so one server cannot exhaust the service for another.
- Your entitlements, meaning plan, languages, rate limit and expiry, are enforced on our side rather than the plugin's, so a modified client cannot widen its own access.
- Revocation is fast and the number is knowable. A revoked licence stops being accepted within about 30 seconds. A token replaced by a routine renewal keeps working for up to 330 seconds, being a 300 second handover window plus up to 30 seconds of cache, so that a server does not go dark mid-renewal.
We do not currently hold an external security certification, and we would rather say that than imply one.
8. Sub-processors
The full list is in Annexe B, including the ones that have nothing to do with audio, so that this document and our privacy policy cannot be read as naming different sets.
The short version: in the audio path there is one, our network provider, because the traffic crosses their infrastructure. Speech recognition itself runs on hardware we operate in the Czech Republic, inside the EU, so the processing needs no transfer mechanism.
We will tell you before adding a sub-processor that touches audio, and you may object by cancelling.
9. Your players' rights, and the opt-out that does not exist yet
For a player who has never been flagged, we hold nothing at all, so there is nothing to export or erase. For a player who has, we hold the moderation records in section 5.2, keyed to their Minecraft UUID, subject to the status note there.
If a player asks you to exercise a right under the GDPR, most of the answer is on your side: the data lives on your server, in your logs, in your Discord if you set a webhook, and in whatever moderation records you keep. Where our records are part of the answer, tell us the UUID and the server and we will export or erase them, within thirty days and at no charge. We will help you answer and we will not obstruct a request.
Erasing a record does not undo a mute your server applied, and it does not oblige you to. Whether a moderation decision stands is yours to judge as controller.
Telling your players. The plugin shows a message to each player when they join, while a transcribing mode is active. It fires on join only, so anybody already connected when you enable moderation is not shown it, and you should announce it yourself the first time. Do not remove it, and correct its wording if you have changed what the product does, for instance by switching on the setting in section 4.
Opting out, stated plainly.
VoiceSniffer does not currently give a player any way to opt out of
transcription by themselves. There is a setting,
consent.opt-out-mode, which decides what happens to a player who is
marked as opted out, and it defaults to switching that player's voice chat off
rather than letting them speak unmoderated. What does not exist is any command, menu
or permission by which a player can become marked. The setting describes a
consequence that nothing currently causes.
We are telling you this here rather than letting you discover it, because if you have told your players they can opt out then that statement is currently untrue and it is your statement, not ours. Until the plugin ships a mechanism, an opt-out on your server has to be something your staff action by hand: take the request however you normally take requests, and act on it. We will update this section when there is a command, and we will not describe one before there is.
10. Children
Minecraft servers have children on them. That is the point of this product and pretending otherwise would be dishonest, so this agreement assumes it rather than disclaiming it.
We do not knowingly collect identifying information about anybody of any age. We are never told a name, an email, an age or an IP address, and we hold no audio. What we hold about a flagged player is a Minecraft UUID and words they said.
You are responsible for whatever consent or parental authorisation your jurisdiction requires from the people on your server, and for age gating if you operate somewhere that demands it. Where your players are children, the bar is higher on all of it: on the notice being understandable to them, on the lawful basis being sound, and in particular on section 12.
11. Breach
If we become aware of a personal-data breach affecting your processing, we will tell you without undue delay and in any case within 48 hours, with what we know and what we are doing about it. You remain responsible for notifying your supervisory authority where that is required.
12. Automated decisions
VoiceSniffer can act on a speech recognition result without a human being involved, and as shipped it does. The default configuration warns on a severity 1 match and mutes the player for five minutes on severity 2 and for an hour on severity 3. Nothing in the product asks a person to confirm that, and nothing in the product offers the player a way to contest it.
The input to that decision is a transcript that is wrong about one word in seventeen in English under good conditions, one in ten in Czech, and worse than either with background noise, a cheap microphone, or an accent the model has heard little of. On the other supported languages we have published no figure. The person it acts on is, on most Minecraft servers, a child.
Nothing gates the decision on confidence. A confidence value is
computed and returned in the verdict, but no code consults it before an action is
applied. A rule that matches a word the model half-heard produces the same mute as
one that matches a word it heard perfectly. There is also no confirmation step, no
appeal route and no dispute workflow anywhere in the product. The only reversal that
exists is a member of your staff running /vs unmute.
Under the GDPR, a decision taken about a person by automated means, which produces a real effect on them, carries obligations: telling people it happens, and giving them a way to obtain human review. Those obligations sit with you, because you are the controller and you configured the actions. Whether they are engaged on your particular server depends on facts about your deployment that we do not have, and we are not in a position to give you legal advice about it. What we can do is tell you exactly what the software does, which is the paragraph above, and give you the controls.
As processor, what we do here is narrow and worth stating: we return a transcript, the rules that matched and a severity. We do not decide anything. The mute is applied by the plugin on your server, according to a configuration you chose. We never see whether it was applied, except where your server chooses to tell us.
What we recommend:
- Set every severity to
warnand have your staff act on the alert. That removes the automated decision entirely and is the configuration we would run. - If you keep automatic mutes, say so in your join message and your rules, and give players a named route to a human who can lift one.
- Do not attach anything irreversible to an automated flag. Bans, public call-outs, and reports to anyone else should have a person behind them.
- Treat a transcript as a reason to look, never as proof of what was said.
13. Deletion, audit, and the end of the agreement
There is no audio to delete on termination, because none was ever kept. The moderation records in section 5.2 are covered by the same retention mechanism that applies while the agreement is live, so ending your subscription needs no request from you and no action from us. Ask and we will delete them sooner. The per-server counters are numbers with no personal data in them and are kept.
Account and billing records are covered by the privacy policy. You may ask us for the information you need to demonstrate compliance and we will provide it.
This agreement lasts as long as your cloud subscription. Where it conflicts with our terms of service, this agreement wins on anything concerning personal data.
Annexe A. The processing, in the Article 28(3) shape
| Subject matter | Speech recognition and rule matching on voice chat audio, so that the controller can moderate its server. |
| Duration | The life of the controller's cloud subscription. |
| Nature and purpose | Decode, transcribe, match against moderation rules, return a verdict. Record flagged utterances so the controller can review them. |
| Types of personal data | Voice audio, transient. Minecraft UUID. The transcript of a flagged utterance and the matched fragment of it. No name, email, age, IP address or position is received. |
| Special category data | Not sought and not inferred. We do not do voice biometrics, speaker identification, emotion detection or any other profiling of the speaker. Speech can nevertheless reveal anything the speaker chooses to say, which is a reason to keep the retention short rather than a claim that it cannot happen. |
| Categories of data subject | Players on the controller's Minecraft servers who use voice chat. Commonly children. |
| Location of processing | Hardware operated by Sniffer Studio in the Czech Republic. |
Annexe B. Sub-processors, and everyone else in the picture
Split by what they touch, because "sub-processor" answers a different question depending on which processing is being asked about.
| Who | What they touch | Status |
|---|---|---|
| Cloudflare, Inc. | TLS termination and network transport for *.snifferstudio.net, including the tunnel that carries audio to the processor. Audio passes through. It is not stored there. |
Sub-processor, and the only one in the audio path. Operates globally and relies on the EU Standard Contractual Clauses. |
| Discord, Inc. | Sign-in for your dashboard and panel account, so your Discord username, user ID and the email on that account. Nothing to do with voice data. | Sub-processor for account data, not for your players' data. Separately, see section 6 for what happens if you point a webhook at Discord: that transfer is yours. |
| BuiltByBit | Sales and payment. They are the seller on almost every licence. We never see card numbers. | Independent controller for the purchase, not our sub-processor for it. |
| Plausible Analytics | Page views on our websites, including this page. Self-hosted by us on our own machine, sets no cookies, stores no IP address, builds no profile. | Not a sub-processor. It is our own software on our own hardware, and it never touches voice data or your account data. |
| Speech recognition | The model that transcribes your players' audio. | No third party at all. It runs on our hardware. There is no OpenAI, Google, Amazon or other speech API anywhere in this system. |
| GitHub and Hugging Face | The one-time download of the speech recognition model files, at first start. Outbound only. | Not a sub-processor. No audio, no transcript and no personal data of any kind is sent to either. They are where the model is fetched from, nothing more. |
Contact
Data protection questions and requests: [email protected], or Discord for anything less formal.