APPLIED AI / SYSTEM CASE STUDY

Synthetic Data Platform

Train AI on worlds you can control.

A completed simulation and data-generation system that turns virtual scenes into photorealistic images, structured annotations and training data for machine learning.

Labelled keypoints with world coordinates in a virtual scene

Body keypoints and world-space coordinates: visual content paired with machine-readable scene information.

01 / SIMULATE

Build the scene

Geometry, actors, movement and sensor configuration.

02 / CAPTURE

Generate the data

Camera images, segmentation, motion and spatial information.

03 / TRAIN

Connect the model

Image-linked annotations for external learning workflows.

04 / EVALUATE

Close the loop

Model predictions and control signals returned to the simulation.

THE CHALLENGE

Useful training data needs more than images.

Real-world recordings capture what a camera sees, but not every object’s exact position, identity, motion or visibility. Creating those labels afterwards takes additional work, while unusual situations can be difficult to capture deliberately.

The platform generates visual data and its annotations from the same virtual scene. Camera configuration, actor state and keypoints accompany the rendered frame, providing an explicit connection between pixels and the underlying 3D world.

The result is a controllable source of synthetic training data for perception, pose estimation, motion analysis and autonomous-system development.

WHAT WAS BUILT

One scene. Multiple learning signals.

The system combines real-time rendering, scene annotation and configurable capture. Each output describes a different aspect of the same environment.

Photorealistic capture

Virtual cameras render the scene from configured viewpoints. Position, orientation, field of view, sensor size, focus and exposure are recorded alongside the images.

Keypoints & actor metadata

Named points on people, vehicles and other objects expose 3D transforms, image-space positions, speed, velocity and visibility. Labels can describe body joints, facial landmarks or object-specific features.

Segmentation & masks

Class-level segmentation groups objects such as pedestrians, vehicles and buildings. Instance-level segmentation isolates individual actors. EXR outputs carry segmentation and mask information.

Optical flow

Motion between frames is represented as image-space displacement. This supports work on movement estimation, tracking and temporal scene understanding.

Stereo disparity

Corresponding views describe the displacement of scene features between cameras, providing a separate signal for stereo vision and spatial perception.

Structured JSON

Each rendered camera frame has a corresponding metadata file containing camera parameters, objects and keypoints. Visual outputs remain associated with their scene information.

ANNOTATION IN PRACTICE

Know where a point is—even when it is hidden.

Keypoints are attached to 3D models and projected into the camera image. Visibility information distinguishes points seen by the camera from those hidden by another object or by the model itself.

Virtual character with body and face keypoints

Named landmarks on a virtual character provide a consistent basis for pose and movement data.

World space and image space

The same point is described relative to its model, in the world and in the rendered frame. The annotations connect a visual observation with the geometry that produced it.

Occlusion as data

A hidden landmark still has a position in the simulation. Recording its visibility state allows a learning workflow to distinguish missing visual evidence from a missing object.

The annotation system also applies to vehicles, signs, buildings and other scene elements; it is not limited to human characters.

Visible and occluded keypoints on a partially hidden virtual character

Visible and occluded landmarks on a partially hidden character.

VISUAL OUTPUTS

Different representations of the environment.

Virtual road scene segmented into object classes

Semantic segmentation separates scene categories for pixel-level interpretation.

Colour-coded motion visualization for a road scene

Optical-flow visualization represents apparent motion across the image.

Grayscale disparity visualization of a street scene

Disparity visualization provides spatial information complementary to the colour image and object annotations.

ARCHITECTURE

From a virtual world to a learning pipeline.

SIMULATION LAYER

Scene & behaviour

Unreal Engine and C++ provide the rendering and simulation foundation. Scene content, actors, traffic behaviour and camera rigs determine what the system observes.

DATA LAYER

Capture & annotation

Rendering and metadata generation produce the camera outputs, segmentation, keypoints and motion information. Each camera exports its own frame-associated data.

MODEL INTERFACE

Training & feedback

External learning systems consume the generated data. The control interface also accepts a model’s outputs and applies them to the virtual environment.

Data path: scene state → sensor capture → images + annotations → learning system → simulated response

The architecture separates data generation from the learning model. It can supply an existing training pipeline or support the complete generate–train–evaluate cycle. The scene, annotation scheme and interface can be adapted to the target application.

APPLICATION / AUTONOMOUS DRIVING

A virtual proving ground for perception and control.

Autonomous driving brings the platform’s capabilities together in one environment: roads, intersections, pedestrians, vehicles and signals provide the scene; virtual sensors observe it; annotations describe it; and a controller acts within it.

Traffic and exceptional situations

The traffic system includes motor vehicles, bicycles, pedestrians, public transport, parking and signalling. Roadworks, obstructions, accidents and emergency vehicles introduce situations that change normal traffic behaviour.

Sensors, navigation and sound

Configurable camera, radar and lidar arrangements represent the vehicle’s sensing setup. Geolocation and navigation supply route information, while acoustic events such as sirens and horns add another source of environmental input.

Annotated vehicles and pedestrians at a simulated urban intersection

An annotated urban environment connects vehicles, pedestrians and scene geometry with the training data.

A complete feedback cycle

A conventional controller drives the virtual vehicle to generate initial data. A deep-learning system consumes the sensor and vehicle information, then returns steering, acceleration, braking and gear values through the control interface. Its behaviour can be observed in the same virtual environment and captured for further analysis.

The platform also accepts externally developed models, allowing perception and control systems to be evaluated against the configured scenarios.

BEYOND DRIVING

The data-generation system is reusable.

Human pose & motion

Character landmarks, transforms and visibility support training data for body tracking, movement interpretation and prediction.

Object & scene perception

Labels, masks and instance information support detection and segmentation across configurable environments.

Visual behaviour analysis

Procedural movement and controlled interactions provide sequences for studying how people and objects move through a scene.

The shared foundation is a configurable virtual environment with explicit annotations. Changing the models, labels, camera setup and behaviours adapts the data to a different learning task.

INTEGRATION

Synthetic data designed around the model.

A useful dataset starts with the learning task: what the model observes, which labels it needs and how its output will be used. The platform brings scene construction, capture and annotation into one system so those requirements can be addressed together.

Core technologies: Unreal Engine · C++ · photorealistic rendering · JSON metadata · EXR segmentation · keypoint annotation · optical flow · stereo disparity

Scroll to Top