desert-ant-core
Add on-device speech recognition, audio cleanup, PII redaction, emoji suggestion, and more to your app, with SDKs for Swift, Kotlin, and JavaScript. The models run on the device, so text, audio, and images never leave it.
import Emo
import Redact
let suggestions = try await Emo().suggestions(for: "Pay my bills") // 💰 💳 🧾
let clean = try await Redact().redaction(of: "Email Anna at anna@example.hu.")
// Email [GIVEN_NAME_1] at [EMAIL_1].
- Models
- Swift
- Android
- JavaScript and TypeScript
- Command line
- Model downloads and caching
- Platform support
- License
Models
| Model | What it does | Platform | Docs |
|---|---|---|---|
| Align | Word-timestamp refinement for any transcript, on device. | Apple · Linux · Windows · Node | SDK Model |
| Clear | On-device speech enhancement: denoise, dereverb, and loudness-normalize. | Apple · Android · Linux · Windows · Web · Node | SDK Model |
| Clips | Short clips and highlights from talking video and audio: podcasts, interviews, meetings. On-device. | Apple · Linux · Windows | SDK Model |
| Ear | On-device spoken language identification across 99 languages. | Apple · Android · Linux · Windows · Web · Node | SDK Model |
| Emo | Multilingual on-device emoji suggestion. | Apple · Android · Linux · Windows · Web · Node | SDK Model |
| Gist | Multilingual on-device content topic tagging across a 36-topic taxonomy. | Apple · Android · Linux · Windows · Web · Node | SDK Model |
| Moderator | On-device NSFW image detection, trained only on licensed and synthetic data. | Apple · Android · Linux · Windows · Web · Node | SDK Model |
| Redact | Multilingual on-device PII detection and redaction. | Apple · Android · Linux · Windows · Web · Node | SDK Model |
| Shapes | On-device single-stroke shape recognition. | Apple · Android · Linux · Windows · Web · Node | SDK Model |
| Title | On-device titles and descriptions: a short factual title and a one- to two-sentence description for any passage of text. | Apple | SDK Model |
| Tongue | On-device language identification for short text across 84 languages. | Apple · Android · Linux · Windows · Web · Node | SDK Model |
| Uhm | On-device filler-word detection: frame-precise "uh"/"um"/"hmm" spans. | Apple | SDK Model |
| Voz | On-device speech recognition: transcripts with word-level timestamps, 25 languages. | Apple · Windows · Web · Node | SDK Model |
In closed beta
Weights exist and the models work, but no SDK ships them yet, so there is nothing to install today. Ask us if you want early access.
| Model | What it does | Docs |
|---|---|---|
| Eye | On-device frame scoring: which shot to keep from a burst or a clip. | Model |
| Face | On-device face matching across a photo library or through a video. | Model |
| Schemer | On-device structured extraction into a caller-supplied JSON schema. | Model |
| Toxic | On-device hate-speech triage for European languages. | Model |
| Who | On-device speaker labeling: per-person turns with timestamps. | Model |
The weights for every model are on Hugging Face.
Swift
On Apple platforms, the Swift package requires iOS 18+, macOS 15+, tvOS 18+, visionOS 2+, and Swift 6.2+ (Xcode 26). An older OS can't load the Core ML models. On Windows, the Swift package requires x64 and Swift 6.2+.
Add the package with Swift Package Manager:
.package(url: "https://github.com/Desert-Ant-Labs/desert-ant-core.git", from: "3.5.0")
Then add a product for each model you use, such as Emo. Your app includes only the models you add.
Android
The Android SDK requires API 24+ and runs on arm64-v8a and x86_64. All models share one copy of LiteRT from ai.desertant:core, so each model you add brings only its own native library.
// settings.gradle.kts
dependencyResolutionManagement {
repositories {
google()
mavenCentral()
}
}
// build.gradle.kts
dependencies {
implementation("ai.desertant:emo:3.5.0")
implementation("ai.desertant:redact:3.5.0")
implementation("ai.desertant:clear:3.5.0")
implementation("ai.desertant:tongue:3.5.0")
}
Tongue has no native code, so you can also use Tongue outside Android, on any JVM 17+.
JavaScript and TypeScript
Each model is its own package, so install the ones you use:
# Browser (WebAssembly + LiteRT.js):
npm i @desert-ant-labs/emo @litertjs/core
# Server-side inference in Node (prebuilt native core, no extra install):
npm i @desert-ant-labs/emo
# Tongue is pure JavaScript, no wasm, no LiteRT.js, no native core:
npm i @desert-ant-labs/tongue
The default import is the browser build. The browser build has no native dependencies, so Next.js, Remix, SvelteKit, and Nuxt can bundle it for every target, including the server-side rendering pass. For inference in plain Node, import the /native subpath. /native ships prebuilt for linux-x64, linux-arm64, and darwin-arm64. Voz has no /native subpath. In Node, pass onnxruntime-node to Voz's load().
Align runs only in Node. You can still import @desert-ant-labs/align anywhere, including the server-side rendering pass. Calling load() from that import throws an error that tells you to import @desert-ant-labs/align/native instead.
Command line
Transcribe a recording, cut clips, clean up audio, or redact text from the terminal. The Desert Ant CLI runs on macOS (Apple silicon) and Linux:
curl -fsSL https://raw.githubusercontent.com/Desert-Ant-Labs/desert-ant-cli/main/install.sh | sh
You can also install the CLI with brew install desert-ant-labs/tap/desertant or with mise. Then run a model:
$ da redact "Email Anna at anna@example.hu or call 555-0100"
Email [GIVEN_NAME_1] at [EMAIL_1] or call [PHONE_1]
Commands pass JSON to each other, so you can transcribe a video once and cut clips from that transcript.
desertant setup writes a skill so coding agents on your machine, such as Claude Code and Codex, can use the CLI.
Model downloads and caching
The SDK downloads each model from Hugging Face on first use. Each SDK version pins one model revision, so your app gets the same model until you update the SDK. The SDK verifies every download before using it. Tongue's 2MB model ships inside each package, so Tongue never downloads anything.
- By default, the SDK keeps the files in the platform cache directory and reuses them across launches.
- Pass
directoryto keep the model in your own folder. If the model files are already in the folder, the SDK uses them and downloads nothing. Otherwise the SDK downloads the model into the folder. - On the web, you can serve the files yourself and pass
modelBaseUrl.
isDownloaded() tells you whether a model works with no network. download() fetches the model ahead of time and reports progress.
Faster downloads on Apple platforms (the Xet trait)
Our Hugging Face repos use Xet storage, which splits each file into deduplicated chunks that download in parallel instead of as one stream. On Apple platforms, the Swift package can download through Xet when you turn on a package trait:
.package(url: "https://github.com/Desert-Ant-Labs/desert-ant-core.git", from: "3.5.0",
traits: ["Xet"])
With the trait on, the same files land in the same cache, and the SDK verifies them the same way. If a file isn't on Xet, or the Xet download fails, the SDK falls back to the ordinary HTTPS download. The trait adds swift-xet and SwiftNIO to your dependencies. Every other platform and SDK downloads over plain HTTPS.
Set HF_TOKEN in the environment for a gated or private repo. Public models need no token.
Offline and airgapped
To run a model offline, download its files and point the SDK at the folder. The SDK then needs no network.
- Find the model's
Catalog.swiftinSources/, at the tag you build against. The catalog lists the Hugging Face repo, the revision, and the files for each platform. - Download those files with the Hugging Face CLI. For example, Redact on Apple platforms:
On Android, Linux, Windows, and the web, the catalog listshf download desert-ant-labs/redact --revision v0.4.0 \ --include "redact.mlmodelc/*" redact_tokenizer.bin labels.json \ --local-dir redactredact.tflitein place of the.mlmodelcfolder. - Point the SDK at the folder:
- Swift:
Redact(directory: path) - Kotlin:
Redact(context, directory = path) - Node:
Redact.load({ directory: path }) - Browser: serve the files yourself and pass
modelBaseUrltoRedact.load().
- Swift:
For a reproducible build, use the revision's commit sha in place of the tag.
AWS Lambda on arm64
Lambda doesn't mount /sys/devices/system/cpu. On arm64, the CPU backend reads that directory to count cores. Without the directory, the library loads but inference fails on every call. Preload the shim that ships alongside the native library:
LD_PRELOAD=/var/task/node_modules/@desert-ant-labs/redact/native/linux-arm64/libdalcpushim.so
Set LD_PRELOAD as a function environment variable. Correct the path if the package lives in a layer (/opt/nodejs/node_modules/...). The shim answers reads of the two files the CPU backend needs there, possible and present, and passes every other call through. The dynamic linker has to insert the shim before libc, so the SDK can't set it up for you.
On x86_64 you don't need the shim, because the core count comes from a CPU instruction instead of sysfs.
Platform support
| Platform | Runtime | Requirements |
|---|---|---|
| iOS, macOS, tvOS, visionOS | Core ML | iOS 18+, macOS 15+, tvOS 18+, visionOS 2+, Swift 6.2+ |
| Android | LiteRT | API 24+, arm64-v8a and x86_64 |
| Windows | LiteRT | x64, Swift 6.2+ |
| Browser | WebAssembly + LiteRT.js | any browser with WebAssembly; @litertjs/core |
| Node | prebuilt native core | linux-x64, linux-arm64, darwin-arm64 |
Voz uses ONNX Runtime outside Apple platforms. On Windows, Voz runs on the GPU through ONNX Runtime and DirectML. In the browser, Voz runs on WebGPU through onnxruntime-web, and uses WebNN where the browser supports it. In Node, Voz runs on the CPU through onnxruntime-node.
License
Desert Ant Labs Source-Available License. Most apps can use the SDK for free. At scale, you need a commercial license. The link has the full terms. For licensing, email licensing@desertant.com. THIRD_PARTY_NOTICES.md lists the third-party components.
