TrackOSC for iOS · TrackOSC Sender for macOS · TrackOSC Receiver for macOS
TrackOSC Recorder · TrackOSC Speaker · TrackOSC Router · TrackOSC Colours · TrackOSC Particles · TrackOSC Text · TrackOSC Synth · TrackOSC Costumes · TrackOSC 3D Costumes – nine more macOS apps that do things with the stream
Live camera → Apple Vision tracking → OSC. TrackOSC streams (almost) all of Apple's Vision framework detection results – body poses in 2D and 3D, hand poses, face landmarks, text, animals (boxes and skeletons), people, and barcodes – over the network as OSC (OpenSoundControl) messages, from an iPhone or a Mac, and visualises them in a companion receiver app.
TrackOSC (formerly Poseiosc) is a native-Swift successor to VisionOSC by LingDong- (itself a successor to PoseOSC) and speaks exactly the same OSC wire format, so existing VisionOSC/PoseOSC receivers (Processing, TouchDesigner, Max/MSP, openFrameworks…) work unchanged.
Nine apps plus open receiver examples for eight creative-coding environments:
- TrackOSC for iOS (iOS 18+, SwiftUI): live camera → Vision → OSC over UDP, with on-screen tracking overlays, per-detector toggles, front/back camera switch, portrait/landscape support with an orientation lock for mounted rigs, and Bonjour discovery of receivers.
- TrackOSC Sender (macOS 15+, SwiftUI): the same tracking pipeline running on a Mac camera – built-in, external webcam, or an iPhone via Continuity Camera – with a camera picker and a rig-rotation setting for cameras mounted sideways.
- TrackOSC Receiver (macOS 15+, SwiftUI): listens on UDP (default port 9527), draws skeletons/landmarks/boxes with coordinate guides, switches to an orbitable 3D view for the 3D body poses, shows per-address message rates and a log, and advertises itself on the local network so senders can find it.
- TrackOSC Recorder (macOS 15+): records the stream to a
.trackoscfile – every datagram exactly as it arrived – and plays recordings back to any receiver at any speed, looping and scrubbing. Develop and demo everything else without a camera. - TrackOSC Speaker (macOS 15+): reads the stream aloud with Apple's speech synthesis – "A person appeared. A hand appeared.", recognised text and codes, and a periodic summary of where everyone is – with every voice and every utterance option exposed.
- TrackOSC Router (macOS 15+): turns tracking events and values into MIDI notes and control changes, Shortcuts, key presses and HTTP requests, by rules you edit in the app – so a raised hand can start a song, a QR code can run an automation, and a nose can turn a knob.
- TrackOSC Colours (macOS 15+, Metal): gradients, colour fields and patterns driven by the stream – fourteen modes from glowing skeletons and Voronoi cells to heat maps and auroras, twelve palettes, nine presets, keyboard control, and a full-screen stage for walls and stages.
- TrackOSC Particles (macOS 15+, Metal): up to a hundred thousand physics particles driven by the stream – attraction, repulsion, orbits, sparks from fast joints, fire along the bones, ghosts of where people were, long exposures, flow fields, constellations, rain and snow that break on the body, a body made of dust, and fountains from the hands.
- TrackOSC Text (macOS 15+, Metal): kinetic typography from the text the sender reads, the codes it scans and your own words – letters that fall and get knocked about, words along skeletons and outlines, word clouds, orbits, scatter, a typewriter, a marquee, box labels and letter rain – in a choice of typefaces.
- TrackOSC Synth (macOS 15+, AVAudioEngine): a 303-and-808-flavoured synth and drum machine played by the stream – an acid bass with a ladder filter, eight analogue-model drums, a sixteen-step sequencer with eight patterns, swing and a conductor mode, and mappings from noses, wrists, hands, faces and presence to every knob and trigger, with MIDI out.
- TrackOSC Costumes (macOS 15+): dresses tracked bodies, faces and
hands in SVG costumes – one file per costume with named layers
(
bone:upperArm:left,head,face:mouth,hand:index…) drawn in Illustrator, Inkscape, Figma or Affinity; three bundled, a folder of your own that reloads as you save, several people at once, mirroring when someone turns their back, and recording to .mp4. - TrackOSC 3D Costumes (macOS 15+, RealityKit): dresses the 3D body pose in a rigged USDZ model on Apple's motion-capture skeleton (the rig ARKit drives, so the Biped Robot and anything made for it works), a folder of parts named per bone, or the built-in mannequin and blocks, with an orbit camera, a settling floor and several people at once.
- Receiver examples for Processing, Python, p5.js, TouchDesigner, Max/MSP, Pure Data, openFrameworks and SuperCollider: each is a complete, hackable receiver of every TrackOSC message – the same drawing (or a sonification) as the native receiver – so you can start making software on the platform you already use, without touching the Apple stack.
A note on the receiver examples. Only the Processing, Python and p5.js examples have been run end-to-end by the maintainer; the TouchDesigner parser is unit-tested but its network recipe hasn't been built in TouchDesigner, and the Max/MSP, Pure Data, openFrameworks and SuperCollider examples were written from those platforms' documentation and checked mechanically (valid patch files, consistent wiring) but never opened in the tool itself, because none of them is installed on the development machine. Treat those as careful first drafts: they follow the same parsing pattern as the verified ones, but expect to fix small things. If you try one, please open an issue or pull request saying which version you used and what changed – that is the most useful contribution this repository can get right now. The per-platform status is tabulated in
Examples/README.md.
The Mac apps are downloadable, notarised builds; the iOS app is free on the App Store (or built from source with your own developer account); the Processing sketch just needs the free Processing editor.
The macOS sender tracking body, hand, and face at 30 fps, streaming OSC to 127.0.0.1:9527.
The macOS receiver drawing the same scene from the OSC stream alone – with per-address message rates, camera info, and a live log.
TrackOSC Speaker reading the stream aloud: what is being said, large enough to read across a room, with the spoken words highlighted.
TrackOSC Router turning appearances into MIDI notes, the nose position into a control change, recognised text into a log line and hand counts into HTTP requests.
TrackOSC Recorder playing a .trackosc file back to the receiver on port 9527 while itself listening on 9528.
TrackOSC Colours in its Body Hue mode, with the mode, parameters and palette controls beside the stage.
All fourteen Colours modes from one synthetic figure (the app's own build check renders this).
The twelve Particles modes (Sparks is empty here because the synthetic figure never moves fast).
The ten Text modes.
TrackOSC Synth: bass and mix, drums, steps and mappings, with the step ring, level glow and drum flashes on the stage.
TrackOSC Costumes: the bundled robot, skeleton and Template on the same tracked figure.
TrackOSC 3D Costumes: the mannequin, the blocks and the rigged Blocky model on the same 3D pose.
Left: the iOS sender. Right: the open Processing (oscP5) receiver sketch drawing the same wire format – no Apple stack required.
All downloads are signed and notarised – no Gatekeeper hoops.
- Mac receiver: download
TrackOSCReceiver-<version>-macOS.zipfrom the Releases page, unzip, and open. Allow the Local Network prompt on first launch. - Mac sender: download
TrackOSCSender-<version>-macOS.zipfrom the same Releases page – the full tracking pipeline running on a Mac camera (built-in, external webcam, or your iPhone via Continuity Camera). Allow the Camera and Local Network prompts. Pick the camera, a rig rotation (for cameras mounted sideways), and the destination in its settings – receivers on the network appear automatically. To try everything on one Mac, run sender and receiver together and send to127.0.0.1. - Mac Recorder, Speaker, Router, Colours, Particles and Text:
TrackOSC<Name>-<version>-macOS.zipfor each, from the same Releases page. Each listens on 9527 like the receiver, so a sender that already works with the receiver works with them unchanged; when 9527 is taken (say the receiver is running), the newcomer takes the next free port and tells you – see Running several apps at once. - iPhone sender: get TrackOSC on the App Store (free, iOS 18+).
- Receiver examples: no Apple anything required – open the
Processing sketch
in Processing with the oscP5 library, run
the Python or p5.js receiver, or
pick TouchDesigner, Max/MSP, Pure Data, openFrameworks or SuperCollider
from
Examples/. See Receiver examples for details.
Everything below is only needed if you want to build from source.
- Xcode 16 or newer (with the iOS 18 and macOS 15 SDKs)
- A Mac running macOS 15 (Sequoia) or newer
- For the iOS sender: an iPhone running iOS 18 or newer
- An Apple developer account (the free tier works)
- Sender and receiver devices on the same Wi-Fi network (guest/hotel networks often block device-to-device traffic – see Troubleshooting)
All dependencies are Swift Packages resolved automatically by Xcode
(swift-osc and the local
PoseioscShared package). Nothing else to install.
- Open
TrackOSC.xcodeprojin Xcode. - Select the TrackOSCReceiver scheme, destination My Mac (the TrackOSCRecorder, TrackOSCSpeaker and TrackOSCRouter schemes build the other three the same way).
- Signing: Xcode may ask you to pick a team – go to the target's Signing & Capabilities tab and select your team (personal is fine).
- Run (⌘R).
- First launch: macOS asks for Local Network permission – allow it, or the receiver can't be discovered (and on some setups can't receive at all). If the macOS firewall prompts about incoming connections, allow those too.
The toolbar shows the listening port (default 9527) and the Bonjour name it's advertising. You can change the port and press Restart.
Same as the receiver, with the TrackOSCSenderMac scheme. First launch
asks for Camera and Local Network permission – both are needed. In
its settings (gear icon): pick a camera (external webcams and iPhones via
Continuity Camera appear automatically), set Rig rotation if the camera
is mounted sideways, and choose a destination – discovered receivers are one
click. Send to 127.0.0.1 to feed a receiver on the same Mac.
- In the same project, select the TrackOSCSender scheme and your iPhone as the destination (connect it by cable the first time).
- In the PoseioscSender target's Signing & Capabilities tab, select
your team. (Target and bundle-ID names keep the historical "Poseiosc" –
bundle IDs are welded to App Store Connect and to users' granted
permissions, so they deliberately never changed with the rename.) If
you're building from source rather than installing from the App Store, also
change the bundle identifier prefix
com.joelgethinlewisto something of your own (e.g.com.yourname.trackosc) – either in Signing & Capabilities, or by editingbundleIdPrefixinproject.ymland runningxcodegen generate. - Run (⌘R). With a free developer account the app must be re-signed every 7 days; paid accounts get a year.
- On the iPhone, if the app won't launch: Settings → General → VPN & Device Management → trust your developer certificate.
- First launch prompts: allow Camera, and allow Local Network (needed both for Bonjour discovery and for sending UDP to your Mac). If you decline Local Network by accident: Settings → Privacy & Security → Local Network → enable TrackOSC.
- Start the receiver on a Mac.
- Start a sender (iPhone or Mac). In its settings (gear icon), the receiver should appear under Discovered receivers within a second or two – tap it. (Or type an IP and port manually – senders can also target TouchDesigner, Max/MSP, Processing, etc. on any port.)
- Point the camera at a person: a skeleton appears on the sender's overlay and, live, on the receiver's canvas.
- Toggle detectors with the chips along the bottom (2D Body / 3D Body /
Hand / Face / Face Landmarks / Text / Animal / Animal Pose / Human /
Barcode / Contours / Horizon / Rectangle – the row scrolls sideways on
the iPhone). Face sends the box with head angles and the jawline;
Face Landmarks sends the full 76-point constellation (eyes, pupils,
brows, nose, lips, jaw) as
/faces/arr, drawn feature by feature on both overlays. One Vision request serves both chips. More detectors = lower frame rate; 2D Body + Hand + Face is the comfortable default. 3D Body runs in its own lane at its own, lower rate so it never slows the other detectors; its rate is shown in Settings → Statistics. The status capsule shows destination, transmitted frame size, and processed fps. - Selfie-style previews are mirrored by default (like the Camera app) so they feel natural – but the OSC coordinates sent to receivers are always unmirrored, matching VisionOSC. Turn the mirror off in the sender's settings if you want the screen to match the receiver exactly.
Tracking quality is best when the declared orientation matches how the camera is actually held, because Vision then analyses unrotated frames.
- iPhone – Auto (default): follows the device as you rotate it between portrait and landscape; the transmitted frame dimensions swap accordingly (e.g. 720×1280 ↔ 1280×720).
- iPhone – Portrait / Landscape Left / Landscape Right (Settings → Camera): locks the assumed orientation. Use this when the phone is mounted – on a tripod, clamped sideways, or lying flat – because automatic detection fails when the phone is flat. If a locked landscape preview appears upside down, pick the other landscape option.
- Mac – Rig rotation (0°/90°/180°/270°): Mac cameras don't rotate on their own, so declare how the camera is physically mounted instead.
The current orientation and dimensions are always broadcast in the
/camerainfo OSC message and shown in the sender's status capsule and the
receiver's canvas.
The shared package includes two CLI tools (run from PoseioscShared/):
swift run poseiosc-testsend 127.0.0.1 9527sends synthetic animated frames of all fifteen message types – point it at
the receiver and you should see a walking stick figure, a waving hand, a
face ring with box and jawline, a "HELLO" text box, a "Cat" box, a 3D
figure two metres from the camera (switch the receiver to 3D), a QR
code, a quadruped skeleton, a human box, two drifting outlines, a rocking
horizon and a rectangle in perspective. Add --landscape to send
landscape-oriented frames instead of portrait. Without Xcode,
python3 Examples/Python/trackosc_testsend.py sends the same scene.
swift run poseiosc-testlisten 9527is a headless decoder that prints one line per received message (quit the receiver app first – only one process can bind the port).
With a real scene recorded once by TrackOSC Recorder (or
Examples/Python/trackosc_record.py), play it back to anything instead:
python3 Examples/Python/trackosc_play.py session.trackosc --loopExamples/ holds a complete receiver for every
TrackOSC message in eight environments – FLOSS starting points for modding
and tinkering with no Apple toolchain required. Each parses the twelve
messages up to v1.4 (the native, Processing, Python and p5.js receivers
also draw the three v1.6 ones), draws (or sonifies) them, and shows the same coordinate guides as
the native receiver; joint orders and edge lists are shared via
Examples/SKELETONS.md.
- Processing (oscP5): the 2D reference sketch, plus a P3D sketch drawing the 3D body poses in real 3D.
- Python (python-osc + pygame): a parser package, a window, a headless printer, and a synthetic sender.
- p5.js: a Node bridge (browsers can't receive UDP) and a sketch; the client works in any web page.
- TouchDesigner: OSC In DAT callbacks that fill Table DATs, a Script SOP, and a network recipe.
- Max/MSP: a
[js]parser, a drawing patch, and a sonification. - Pure Data (vanilla): parse and joint-picking abstractions plus a sonification demo.
- openFrameworks (ofxOsc): a reusable C++ parser with 2D and 3D views.
- SuperCollider: OSCdefs, sonification, and a drawing window.
The Processing, Python and p5.js examples were run by the maintainer; the others were written from their platforms' documentation and are waiting for someone with that tool installed to confirm them – pull requests welcome.
Every macOS app in this repository that receives OSC – Receiver, Recorder, Speaker, Router – listens on UDP 9527 by default, the port both senders target, so any one of them works out of the box. Only one process can own a port, so when 9527 is already taken the app that launches next falls forward to 9528, 9529… and shows a banner naming the port it got (and advertises itself on that port, so it still appears in the senders' Discovered receivers list). A port you type in explicitly is never changed behind your back.
To feed one stream to several apps on one Mac, chain them: in the app
that has 9527, open the Forward popover (the ↳ toolbar button) and forward
to 127.0.0.1:9528; that app can forward on to 9529, and so on. Forwarding
re-sends every datagram unchanged, so nothing is lost or re-encoded. A
sender can equally be pointed straight at any of the ports.
The receiver-type apps are built to run in installations and on stage.
Full Screen → Enter Full Screen (⌘⇧F) fills the screen with the stage
alone – the visualiser, the spoken sentence, the rule LEDs – on a flat
background with no title bar, toolbar, controls or cursor. Esc brings
everything back; ⌘⇧H hides or shows the controls without leaving full
screen; the menu also offers Always on Top when windowed and Start in
Full Screen so a Mac that boots into the app shows nothing else, and a
--fullscreen launch argument does the same for one launch:
open -a "TrackOSC Receiver" --args --fullscreenThe background colour is the colour well in the toolbar: black by default, but any colour, and it fills the whole window in full screen.
When more than one sender targets the same port – an iPhone and the Recorder playing a file back, say – the Senders toolbar popover decides who is heard. Latest sender wins (the default) gives the stream to whichever host most recently started sending, so pressing Play in the Recorder takes over from the phone, and the phone gets the stream back a second after playback stops. All senders mixes everything as it arrives (the pre-1.6 behaviour) and Only one host pins a single address; the popover lists the hosts heard in the last minute and the status panel counts what was ignored. Ignored datagrams are invisible to forwarding and recording too.
Records the incoming stream to a .trackosc file and plays one back to any
host and port. Recording is a tap on the raw datagrams, so the file holds
every message exactly as sent – including anything this version doesn't
decode – and playback reproduces the stream byte for byte at 0.25× to 4×,
looping, with a scrubber; after a seek the last /camerainfo is re-sent so
the receiver knows the frame size. Files go to ~/Downloads/TrackOSC Recordings or a folder you choose; Record on launch logs whole
sessions; a file double-clicked in the Finder opens and plays. The same
files are read and written by
trackosc_record.py and trackosc_play.py (standard
library only), and the format is documented in
Examples/RECORDING_FORMAT.md. Record a
minute of a real scene once and every other app – and every receiver
example – can be developed on the train.
Narrates the stream with AVSpeechSynthesizer. Events are debounced so a
flickering detector doesn't produce a flickering commentary: a count has
to hold for 0.4 s to be believed and 0.8 s to be dropped. What gets said:
appearances, departures and count changes per message kind (tick the kinds
you want; body, human and 3D body all mean "a person", so tick one),
recognised text ("I can read: HELLO"), codes ("QR code: github.com, JGL,
TrackOSC" – URLs are read as words, not letters), animal names, and a
summary every n seconds at three levels of detail (counts; nose position
and raised hands; face direction, mouth open, distance and height from the
3D body). Speech exposes every utterance option – voice, rate within
the API's bounds, pitch 0.5–2×, volume, pauses before and after, the
assistive-technology preference, SSML – and how new sentences meet old
ones (latest wins, or a capped queue; interrupt immediately or at a word
boundary). Voices lists every voice installed on the Mac with
language, quality, gender and novelty/personal filters and a preview
button; more voices, including Enhanced and Premium ones, come from System
Settings → Accessibility → Spoken Content. Transcript shows what was
said and can append it to a text file.
Rules turn the stream into actions. A rule is a trigger – a message kind appearing, leaving or changing count; a value rising above or falling below a threshold (with hysteresis); a value entering or leaving a range; a value mapped continuously from an input range to an output range at a chosen rate; or text matching a pattern (equals, contains, regular expression, with a cooldown) – over a source – counts, the nose or any body or hand joint as a 0–1 fraction of the frame, raised hands, face centre, width and angles, mouth openness, distance and height from the 3D body, recognised text, code payloads, animal labels – and an action:
- MIDI note (with velocity and length) or control change, sent from a virtual MIDI source called TrackOSC Router that any DAW, synth or lighting desk can pick as an input, and optionally to a hardware destination too. A continuous rule drives the CC value or note velocity.
- Shortcut: runs a Shortcut by name, with
{value},{text}or{rule}as its input. macOS asks once for permission to control Shortcuts Events; if that is refused, theshortcuts://URL scheme is used instead. - Key press: any key with modifiers, sent to the frontmost app – for remote clickers, games, video players. Needs Accessibility access (the Outputs tab requests it).
- HTTP GET or POST with templated URL and body – lights, OBS, Home Assistant, your own server.
- Log only, for checking a trigger before wiring it up.
The stage shows an LED per rule and the activity feed; Dry run logs what would happen without sending; Sources shows every value live; rules save as JSON in Application Support and can be imported, exported and started from ten presets.
Colour driven by people. Every frame the receiver's latest messages become a tracking scene – people with stable identities and smoothed joints in 0–1 coordinates (mirrored, like a selfie, by default), their hands and faces, recognised text, and two moods, presence and activity – and a Metal shader paints the whole screen from it. Fourteen modes: Body Hue (glowing skeletons, a colour per person), Hand Glow, Joint Stops (every joint a colour stop of one smooth field), Voronoi People, Metaballs, Rings, Stripes (turning with the shoulders), Checkers (warped by whoever stands in it), Memory Wash and Heat Map (which remember where people were), Kaleido Body, Aurora, Face Mood (colour follows the gaze and the mouth) and Palette Sweep. Each has three parameters; twelve palettes are built in and any can be edited; nine preset slots save mode, parameters and palette together. When nobody has been tracked for a while a synthetic figure wanders through so the wall never goes dead (Display → Attract after).
Every mode uses whatever is arriving: cats and dogs from Animal Pose are tracked and drawn exactly like people (their own colours); with Hand and Face Landmarks on, hand skeletons and a face ring join the body in the drawing modes, hand joints and face centres become extra stops, seeds, warps and heat sources in the field modes, and mouth openness and hand openness push the aurora, stripes and metaballs about. A person's head is drawn from their face landmarks when those arrive, otherwise as a circle.
Recording: the Record button (or the V key) writes the rendered
output – not a screen grab – to an H.264 .mp4 in Downloads at the
window's resolution and 60 fps, ready to post; a portrait window gives a
portrait video. Keys on the stage: ←/→ mode, Space randomise, R reset,
1–9 load a preset (⇧1–9 saves), S screenshot, V record, H hide the
controls, F full screen. Frames are fitted (letterboxed) into the window by default; the
modes still fill the screen, only the people's positions are mapped into
the frame area. Adding a mode is one fragment function in
Colours/Shaders/ColoursModes.metal and one catalogue entry in
Colours/ColoursModes.swift; the app's --snapshot-dir <folder>
argument renders every mode to a PNG without a window, which is how the
contact sheet above was made:
open -a "TrackOSC Colours" --args --snapshot-dir ~/Downloads/colours-checkThe same tracking scene, drawn as particles simulated on the CPU and
rendered as additive sprites over a fading background (the Trails
parameter is how much of the last frame survives). Twelve modes: Attract
(every particle races for the nearest joint), Repel (a field parts
around the body), Orbit (moons around each person), Sparks (fast
joints throw sparks that fall), Skeleton Fire (bones burn), Ghost
Parade (where each person was, half a second ago and before), Long
Exposure (joints leave light that lingers), Flow Field (a current that
moving joints stir), Constellation (stars, and the lines that make a
figure of them), Rain / Snow (drops that break on the body), Particle
Body (the body made of dust) and Hand Fountains. Each mode has a
particle count (up to 100,000), size, speed, trails and one parameter of
its own; palettes, presets, keys, recording and the --snapshot-dir check
are the same as Colours. Cats, dogs, hands and faces take part everywhere.
Kinetic typography. The words come from three places: text the sender
reads (/texts/arr), codes it scans (/barcodes/arr), both kept for half a
minute after they were last seen, and your own words typed into the Words
box (one or more per line). Letters are sprites from a glyph atlas built
from the chosen typeface, so they move at the display rate and record like
everything else. Ten modes: Physics Letters (fall, pile up, get knocked
about by the body), Along Skeleton (words run along arms, spine and
legs), Along Contour (words trace /contours/arr outlines, or a jaw,
or the body's outline), Word Cloud, Orbit (letters circle every
joint), Scatter (a resting sentence that fast movements scatter),
Typewriter (recognised text typed out where it was read), Marquee
(words scroll past at the height of each head), Box Labels (text and
codes where they were seen) and Letter Rain. However many words there
are, they repeat to fill whatever they run along; heads are the face
landmarks when they arrive and a circle otherwise, and cats and dogs get
words along their spines, legs, ears and tails, marquee rows and outlines
just like people.
A synth and drum machine in the spirit of the Roland 303, 606 and 808,
played by whoever the sender is tracking. The bass is a saw or pulse
through a four-pole ladder filter with envelope modulation, accent, slide
and overdrive; the drums are analogue models (a swept sine kick, a
two-tone snare with filtered noise, hats from six square waves at the 808's
ratios, two toms, a clap of noise bursts and a cowbell), each with tune,
decay, tone and level. The Steps tab is a sixteen-step sequencer: a bass
row (click to gate a step, drag for the note, ⌥-click for accent, ⇧-click
for slide) and a row per drum (click cycles off, on, accent), eight
patterns, tempo, swing, randomise and clear. Conductor mode stops the
clock and lets a gesture advance the steps instead. The Mapping tab
connects tracking to the instrument: continuous sources (nose across and
down, wrist heights, movement speed, hand openness and spread, hands apart,
3D body height and distance, face yaw, roll and pitch, mouth openness,
number of people, presence, activity) drive any knob, the tempo, the mix or
a MIDI CC through an input range, an output range, a curve, smoothing and
invert, and event sources (a hand raised, a hit, hands together, mouth
opened, a code changing, a person entering or leaving, every beat) fire
drums, bass notes, MIDI notes, pattern changes, step toggles, the conductor
step or play/stop. Three presets to start from: Acid theremin (nose
sweeps the filter, wrists set resonance and decay), Drum conductor
(hits advance the steps, hands play clap and cowbell) and Two-hand
filter (hands apart opens the filter, hand openness sets resonance).
Presets hold the whole instrument, patterns and mappings, and can be
saved, imported and exported as JSON. Everything the synth plays also goes
out of a virtual MIDI source called "TrackOSC Synth" (bass on channel 1,
drums on channel 10 as General MIDI notes, optional clock and start/stop)
and, if you choose one, to a MIDI destination. The stage is black with the
sixteen steps as a ring of lights, the level as a glow, drum hits as
flashes and the tracked figure faintly behind; the audio keeps running with
the controls hidden and in full screen. The DSP lives in the SynthCore
package (cd SynthCore && swift test renders it offline: no NaNs, filter
stability at full resonance, PolyBLEP aliasing, sequencer timing).
Cut-out puppetry: a costume is one SVG file whose layers are named after
the parts they dress, and the app places each layer on the tracked person
at the display rate. Bones (bone:torso, bone:upperArm:left,
bone:shin:right, …) map the layer's art axis (a pivot line you draw,
or the shape's midline) onto the live joint pair, scaling uniformly, only
along the bone (.stretch) or not at all (.fixed); head spans the
ears; face: parts sit on the Face Landmarks clusters (the mouth opens
with the lips) and ride on the head when landmarks are not arriving;
hand: parts follow the Hands detector's fingers, per finger or per
phalanx, and sit at the body's wrist without it. When someone turns their
back the shoulders cross and every layer without .noflip is mirrored;
.front and .back layers show only one way round. Parts whose joints
vanish fade out and hold their last place. Layer names are read from
Inkscape labels, Figma and Illustrator ids (with Illustrator's _x3A_
escapes undone) and Affinity's serif:id; the parser handles paths with
arcs, transforms, <style> classes, style="" and presentation
attributes, and lists what it skipped (gradients, clones, clip paths,
text, images) in the Layers tab. Three costumes are bundled
(robot, skeleton and an annotated Template to copy from), and
Choose Folder… points the app at a folder of your own that reloads
whenever a file is saved, so a drawing app can stay open beside the
stage. With several people tracked, everyone wears the same costume or
each new arrival gets the next one in the library. V records the
stage to an .mp4 in Downloads, M mirrors, K shows the tracked
skeleton, [ and ] change costume. The parser and rig are the
CostumeCore package (cd CostumeCore && swift test runs fixtures in
the shapes each drawing app exports); the recipe for each app is in
Costumes/Resources/Costumes/README.md.
The 3D sibling of Costumes, for /poses3d/arr (turn on 3D Body on
the sender). Three ways to dress the pose. A rigged model (.usdz,
.usd, .usda, .usdc or .reality) whose skeleton follows Apple's
motion-capture rig: the 91 joints named root, hips_joint,
spine_1_joint … left_forearm_joint, right_upLeg_joint and so on, in
Apple's hierarchy, T-posed with +Y up, facing +Z, the left hand along +X
and each joint's +X pointing down its bone. That is the rig ARKit's body
tracking drives, so a model made for it, Apple's own Biped Robot, a
renamed Mixamo rig or anything from Maya, Blender or Cinema 4D via
Reality Converter works here; TrackOSC drives the hips, spine, neck,
head, shoulders, arms, hands, legs and feet by forward kinematics (each
bone is aimed along the live joint pair in its parent's frame, the hips
and spine also take the hip and shoulder lines for yaw and twist, the
model keeps its own bone lengths and is scaled to the person's height,
and only the hips translate), and the Rig tab lists any joints that
are missing. A folder of parts needs no rigging: one model per bone
named torso, head, forearm-left, shin-right and so on, each laid
along its bone by its longest axis and scaled to the bone's length. And
the built-in Mannequin (capsules) and Blocks show something the
moment 3D poses arrive. Blocky.usda, bundled, is a hand-written rigged
model on the exact rig, readable in a text editor. The stage is the
Receiver's 3D stage: metres, a floor that settles under the lowest
ankle, an axis gnomon and camera marker at the origin, drag to orbit
and pinch to zoom; several people each get a costume, the same or
cycling through the library, fading when they leave. The rig maths is
the Costume3DCore package (cd Costume3DCore && swift test), and
swift run costume3d-blocky out.usda writes the example model. The
models README is in Costumes3D/Resources/README.md.
Both senders have a Hide video preview option (the eye button on the main screen, or the toggle in Settings): the camera and OSC output keep running, but the screen shows only the tracking overlay on black – useful on stage or in installations where the raw camera feed shouldn't be visible.
If you have a paid Apple Developer Program membership, you can put the sender on TestFlight so testers (e.g. students) install it from a link – no Xcode needed on their side.
One-time setup:
- Sign in at App Store Connect → Apps → + → New App: platform iOS, a name (e.g. "TrackOSC"), your bundle ID (register it under Certificates, IDs & Profiles or let Xcode's signing pane register it first), any SKU string.
- In Xcode, select the TrackOSCSender scheme with your team set.
Per release:
- Bump
CURRENT_PROJECT_VERSIONinproject.yml(every upload needs a higher build number) and runxcodegen generate– or edit the build number in Xcode's target General tab. - Select destination Any iOS Device (arm64) → Product → Archive.
- In the Organizer window: Distribute App → TestFlight & App Store → Upload (accept the defaults).
- In App Store Connect → your app → TestFlight tab: wait a few minutes
for processing, then either
- add Internal Testers (up to 100 App Store Connect users – instant), or
- create an External group and enable a public link (up to 10,000 testers – ideal for a class; the first build needs a one-off Beta App Review, usually a day or two).
- Testers install the TestFlight app from the App Store and open your public link on the iPhone. Builds expire after 90 days – upload a new one before term ends!
The project already sets ITSAppUsesNonExemptEncryption to false (the app
contains no custom cryptography), so uploads skip the export-compliance
question.
Scripts/release.sh archives every macOS app (receiver, sender,
recorder, speaker, router), signs them with your Developer ID, submits them
all for notarisation in one batch, staples them, and publishes the zips as
one GitHub Release – so end users can download and double-click with no
Gatekeeper friction.
One-time setup:
- A Developer ID Application certificate: Xcode → Settings → Accounts → your team → Manage Certificates → +.
- Notary credentials in your keychain (use an app-specific password):
xcrun notarytool store-credentials poseiosc-notary --apple-id you@example.com --team-id YOURTEAMID --password your-app-specific-password- The GitHub CLI authenticated (
gh auth login).
Then, per release – bump MARKETING_VERSION in project.yml, run
xcodegen generate, commit, and:
POSEIOSC_TEAM_ID=YOURTEAMID Scripts/release.shAdd --dry-run to build/notarise without publishing, --only A,B to
release only some schemes, and --skip-notarize for a local signing test.
App icons are rendered from Design/render_icons.swift by
Scripts/make_appiconsets.sh.
Everything on the wire uses one convention – the same one VisionOSC uses:
(0,0) ──────────► x width
│ ┌───────────────────────────────┐
│ │ │
▼ │ pixels, y down │ height
y │ │
└───────────────────────────(w,h)
- Pixels, not normalised: divide by the frame width/height from the
message header (or
/camerainfo) to normalise. - Origin top-left, y grows downward (screen convention, not math convention).
- Never mirrored. The selfie-mirror option only flips the phone's display; wire coordinates are always the unmirrored scene.
- Dimensions follow orientation: portrait sends 720×1280, landscape
1280×720. Listen to
/camerainfoand your mapping code needs no special-casing:
// Processing: map a TrackOSC point into your sketch window
float sx = x / frameW * width; // frameW/frameH from the message header
float sy = y / frameH * height;(For a full working example – parsing, skeletons, coordinate guides – see
Examples/Processing/TrackOSCReceiver.)
- A keypoint that wasn't detected arrives as
x=0, y=frameHeight, confidence=0– always filter onconfidence == 0.
Byte-compatible with VisionOSC. All messages are sent unbundled over UDP, one per enabled detector per processed frame, including when nothing is detected (header-only). Coordinates are as described above.
/camerainfo (TrackOSC addition; VisionOSC receivers ignore it) –
sent with every processed frame:
| # | Type | Value |
|---|---|---|
| 0 | int32 | frame width in pixels |
| 1 | int32 | frame height in pixels |
| 2 | int32 | orientation: 0 = landscape, 90 = portrait, 180 = landscape flipped, 270 = portrait upside-down |
| 3 | int32 | camera facing: 0 = back, 1 = front |
Every message begins with the same header:
| # | Type | Value |
|---|---|---|
| 0 | int32 | frame width (oriented pixels, e.g. 720) |
| 1 | int32 | frame height (e.g. 1280) |
| 2 | int32 | number of detections n (capped at 32) |
Then, per detection:
/poses/arr – float confidence, then 17 joints × (float x, float y,
float confidence). Joint order (PoseNet order): nose, leftEye, rightEye,
leftEar, rightEar, leftShoulder, rightShoulder, leftElbow, rightElbow,
leftWrist, rightWrist, leftHip, rightHip, leftKnee, rightKnee, leftAnkle,
rightAnkle.
/hands/arr – float confidence, then 21 joints × (x, y, confidence).
Order: wrist; thumb CMC, MP, IP, tip; index MCP, PIP, DIP, tip; middle …;
ring …; pinky … .
/faces/arr – float confidence, then 76 landmark points ×
(x, y, precisionEstimate), in Vision's constellation order.
/texts/arr – float confidence, float left, float top,
float width, float height, string recognised text.
/animals/arr – float confidence, float left, float top,
float width, float height, string label ("Cat" or "Dog").
A joint/point that wasn't detected is sent as x=0, y=frameHeight, confidence=0 (VisionOSC's convention) – filter on confidence == 0.
VisionOSC's /faces/arr carries only the 76 landmark points – no face
boundary. TrackOSC adds two messages (VisionOSC receivers ignore them; the
five messages above are untouched). Both start with the standard
width/height/n header, and both list the same faces in the same order,
so index i in one matches index i in the other. (/faces/arr applies a
stricter landmark check and can, in rare cases, contain fewer faces – don't
assume its indices line up with these.)
/faces/box – per face, 8 floats (fixed stride: face i starts at
argument 3 + i×8):
| Type | Value |
|---|---|
| float | confidence |
| float × 4 | bounding box: left, top, width, height (pixels, origin top-left) |
| float × 3 | head rotation: roll, yaw, pitch in degrees (0 when unavailable) |
/faces/contour – per face: float confidence, int32 m (contour
point count), then m × (float x, float y). The contour is the jawline –
an open polyline from ear to chin to ear; don't close it. m varies by
OS version (typically 17) and is 0 when Vision reports no contour for that
face – always loop on m, never hardcode it.
Four more additive messages, each behind its own detector chip. VisionOSC receivers ignore them; the messages above are untouched. All start with the standard width/height/n header, and every coordinate that is a pixel uses the same convention as everything else (oriented frame, origin top-left, never mirrored).
/poses3d/arr – per pose (fixed stride: pose i starts at argument
3 + i×87; joint j of pose i at 3 + i×87 + 2 + j×5):
| Type | Value |
|---|---|
| float | confidence |
| float | estimated body height in metres |
| 17 × (float x, float y, float z, float px, float py) | per joint: position in metres in Vision's camera-relative space, then the same joint projected into the frame in pixels |
Joint order (Vision's 3D skeleton, root first – not the PoseNet order of
/poses/arr): root, spine, centerShoulder, centerHead, topHead,
leftShoulder, leftElbow, leftWrist, rightShoulder, rightElbow, rightWrist,
leftHip, leftKnee, leftAnkle, rightHip, rightKnee, rightAnkle. All 17 joints
are always present – there is no missing-joint sentinel. The metric space is
right-handed with the camera at the origin, x to the right and y up; z runs
along the camera's axis (the receiver's 3D view and the Processing 3D
sketch each have a single sign constant should your device report it the
other way). The pixel projections let 2D-only receivers draw the 3D
skeleton without any projection maths. On iPhone the heights are measured
using the camera's intrinsics; on a Mac they are estimated from a reference
height. /poses3d/arr is sent at the 3D detector's own rate, typically
lower than the other messages.
/barcodes/arr – per code: float confidence, float left, top,
width, height (axis-aligned box), then 4 × (float x, float y) – the
corners of the code in its own orientation, top-left, top-right,
bottom-right, bottom-left (draw them as a closed quadrilateral) – then
string symbology (QR, EAN13, Code128, DataMatrix, Aztec,
PDF417, MicroQR, …) and string payload (empty when none).
/animalposes/arr – per animal: float confidence, then 25 joints ×
(float x, float y, float confidence), same conventions as
/poses/arr including the missing-joint sentinel. Joint order: nose,
leftEye, rightEye, leftEarTop, leftEarMiddle, leftEarBottom, rightEarTop,
rightEarMiddle, rightEarBottom, neck, leftFrontElbow, leftFrontKnee,
leftFrontPaw, rightFrontElbow, rightFrontKnee, rightFrontPaw,
leftBackElbow, leftBackKnee, leftBackPaw, rightBackElbow, rightBackKnee,
rightBackPaw, tailTop, tailMiddle, tailBottom.
/humans/arr – per person: float confidence, float left, top,
width, height. A whole-body box with no skeleton – cheap, and it counts
people at distances where pose estimation gives up.
Edge lists for drawing all four skeletons, and the colours the apps use,
are in Examples/SKELETONS.md.
Three more additive messages behind three chips at the end of the row,
from Vision's DetectContoursRequest, DetectHorizonRequest and
DetectRectanglesRequest. Same header, same pixel conventions.
/contours/arr – per contour: float confidence, int32 m, then m ×
(float x, float y): a closed outline (draw it with the last point
joined back to the first – unlike /faces/contour, which is open). Vision
returns every edge it finds, nested holes included, walked outline-first;
the sender simplifies each outline and caps a message at 64 contours and
4,000 points so it fits one datagram. Detection runs on a 512-pixel copy of
the frame, dark shapes on a light background.
/horizon – n is 0 or 1; when 1: float confidence, float angle in
degrees, then float x1, y1, x2, y2 – the horizon as a line through the
frame's centre from the left edge to the right edge. A positive angle
raises the right-hand end (one constant in the sender,
horizonPositiveRaisesRight, should a device say otherwise). Point the
camera at a real horizon or a tabletop edge to see it.
/rectangles/arr – per rectangle: float confidence, float left,
top, width, height, then 4 × (float x, float y) corners in the
rectangle's own orientation (top-left, top-right, bottom-right,
bottom-left), exactly like a barcode without the strings. Up to 16 per
frame, at least a tenth of the frame in size, confidence 0.5 or better –
screens, sheets of paper, picture frames, doors.
project.yml XcodeGen spec (source of truth for the Xcode project)
TrackOSC.xcodeproj Generated project (committed; users just open it)
PoseioscShared/ Swift package: wire format codec, models, skeleton
edge lists, coordinate mapping, CLI test tools, tests
SenderCore/ Platform-neutral sender pipeline shared by both
senders (Vision processing, OSC, Bonjour, overlay)
Sender/ iOS sender app shell (camera, rotation, UI)
SenderMac/ macOS sender app shell (camera picker, rig rotation)
ReceiverCore/ Shared by every receiver-type macOS app: UDP listener,
decoding, forwarding, port fall-forward, Bonjour, settings,
full screen, window shell, presence/metrics analysis,
tracking scene (person tracker, history, attract figure),
the RealityKit 3D stage (floor, gnomon, orbit camera)
Receiver/ macOS receiver (2D + 3D visualisers, log)
Recorder/ macOS recorder/player (.trackosc files)
Speaker/ macOS speaker (narration engine, AVSpeech, voice catalogue)
Router/ macOS router (rules, MIDI/Shortcuts/keys/HTTP actions)
VisualCore/ Shared by the visual apps: Metal canvas and renderer (shader
modes plus a sprite/line layer and glyph atlas), recorder,
palettes, parameters, presets, inspector UI
Colours/ macOS colours app (fourteen shader modes)
Particles/ macOS particles app (CPU simulation, twelve behaviours)
Text/ macOS kinetic text app (word pool, letter system, ten behaviours)
SynthCore/ Swift package: DSP (PolyBLEP oscillators, ladder and state-variable
filters, envelopes), 303 bass and 808 drum voices, lock-free
parameter bank and event queues, audio graph and AVAudioEngine
host, step sequencer, mapping logic, MIDI words, presets, tests
Synth/ macOS synth app (knobs, drums, sequencer grid, mapping table, MIDI out)
CostumeCore/ Swift package: SVG-subset parser (XMLParser, path grammar, transforms,
CSS classes), layer-name grammar, 2D rig (bones, head, face, hands,
mirroring, fades), CoreGraphics renderer, fixture tests
Costumes/ macOS costumes app (library with folder bookmarks and hot reload,
Canvas stage, mp4 recorder, bundled costumes and recipe README)
Costume3DCore/ Swift package: Apple's 91-joint motion-capture rig table, T-pose rest,
17-joint retargeting, FK solver, parts maths, USDA writer, tests
Costumes3D/ macOS 3D costumes app (RealityKit stage shared with the Receiver,
rigged/parts/primitive costume entities, library, Blocky.usda)
Examples/ Receiver examples: Processing, Python, p5.js, TouchDesigner,
Max/MSP, Pure Data, openFrameworks, SuperCollider – see
Examples/README.md; Examples/SKELETONS.md is the shared
joint-order/edge-list reference; Examples/RECORDING_FORMAT.md
specifies .trackosc files
Design/ App icon renderer
Scripts/ Notarised-release tooling
PROMPTS_AND_DECISIONS.md Running record of prompts and design decisions
Contributing: source files added/removed? Run brew install xcodegen once,
then xcodegen generate and commit both project.yml and the regenerated
project. Wire-format changes must keep the golden-bytes test in
PoseioscShared/Tests green – that test is the VisionOSC compatibility
contract. Run tests with cd PoseioscShared && swift test. House style: en dashes (U+2013), never em dashes (U+2014), in code, comments and prose alike; Scripts/release.sh refuses to release while any tracked file contains one.
- Receiver never appears in a sender's list – both devices on the same Wi-Fi?
Many guest/campus/hotel networks enable client isolation, which blocks
device-to-device traffic entirely; use a private network or a personal
hotspot. Check Local Network permission on both devices, and that the
Mac's firewall allowed TrackOSC Receiver. Verify the Mac is advertising:
dns-sd -B _osc._udp local. - Discovered receiver won't resolve – enter the Mac's IP manually (System Settings → Wi-Fi → Details → IP address).
- Data sends but nothing draws – confirm the port matches the receiver toolbar; check the receiver's total-message counter is climbing.
- Overlay/receiver skeleton is rotated or flipped – the sender assumes a portrait phone; keep the phone upright. (If it's still wrong on your device, please open an issue naming the iPhone model.)
- Low frame rate – disable Text/Animal (they're the slow ones), good lighting helps every detector.
- Port 9527 already in use – Protokol, OSC DataMonitor, or another receiver may be bound to it; only one process can listen per port.
Two other free tools turn cameras into OSC-speaking people sensors for creative coding; each covers ground TrackOSC doesn't, and vice versa.
- openTSPS – the "Toolkit for Sensing People in Spaces" from the LAB at Rockwell Group. A desktop openFrameworks/OpenCV app that runs blob and person tracking on webcams, Kinect and other depth sensors, or video files, and broadcasts the results as OSC, TUIO, TCP, or JSON over WebSockets. The classic choice for overhead or depth-camera installations that need presence, position and size of people rather than skeletons; it is in maintenance mode (built on openFrameworks 0.8), but still works and is well documented.
- tramontanaCV – by Pierluigi Dalla Rosa, part of the Tramontana platform for prototyping interactive spaces with phones. An iOS app that uses the phone as a sensor, running blob detection or face tracking on its camera and sending the results over Wi-Fi to Processing (and, through Tramontana's libraries, to p5.js, JavaScript and openFrameworks). The closest precedent for TrackOSC's phone-as-tracker idea, with a lighter, blob-oriented output.
TrackOSC's niche next to them: Apple Vision's richer detectors – full body skeletons in 2D and 3D, hand and face landmarks, animal skeletons, text and barcodes – from any iPhone or Mac camera, speaking the VisionOSC wire format so existing receivers keep working. If you need depth-camera tracking of a whole room, reach for openTSPS; if you need the simplest possible blob-from-a-phone, tramontanaCV; if you need skeletons and landmarks, TrackOSC.
Inspired by VisionOSC and PoseOSC by LingDong-, which grew out of ofxFaceTracker by Kyle McDonald. OSC via swift-osc by Steffan Andrews.
MIT – see LICENSE.
TrackOSC exists because Golan Levin suggested building it in the first place – a native, freely available successor to VisionOSC that students could point at any receiver – and then test-drove every release, sending the feedback that shaped the face boundary messages, the Processing example, the screenshots, and the receiver examples for other platforms. Thank you, Golan.













