Technology
Image
Automates data normalization by resizing images to 224x224 via Pillow and transcoding audio into uniform 16kHz mono formats.
This workflow automates the heavy lifting of data preparation for multimodal AI. We use Pillow to force images into a 224x224 pixel square (the standard for ResNet and VGG architectures) while maintaining aspect ratio through smart padding. On the audio side, we leverage FFmpeg to transcode diverse formats into 16kHz mono WAV files: this ensures consistent sample rates for downstream spectrogram generation. It is a no-nonsense approach to cleaning noise and unifying inputs before they hit the training loop.
What builders pair with Image
Projects using both technologies. Select a pairing to see a project.
12 more pairings
Pairing: Next
Multi-Pass Building Defect Detection: Getting a VLM to Find Facade Defects for Visual Inspections
Pairing: AWS
Multi-Pass Building Defect Detection: Getting a VLM to Find Facade Defects for Visual Inspections
Pairing: BAML
Extract anything - from any model
Pairing: Claude Sonnet
Multi-Pass Building Defect Detection: Getting a VLM to Find Facade Defects for Visual Inspections
Pairing: Flash
Building NousyBooks - Orchestrating Low-Latency Multimodal Voice Agents with Gemini Live
Pairing: Gemini-2
Building NousyBooks - Orchestrating Low-Latency Multimodal Voice Agents with Gemini Live
Recent Talks & Demos
Showing 1-4 of 4