TECHPERFORMANCE
AI · Media production

On-premise transcription and AI vision.
For newsrooms and broadcasters.

Whisper for audio/video transcription, ComfyUI for image workflows, vision models (Qwen 2.5-VL, LLaVA) for automatic tagging. All on dedicated EU GPUs, overnight batch integrated into the CMS, no uploads to third-party cloud vendors.

Reference stack

Open models, dedicated GPUs, orchestrated pipelines.

  • → Whisper large-v3 for transcription and diarisation: Italian, EU languages, accuracy comparable to cloud vendors.
  • → ComfyUI as the image workflow engine: editorial thumbnails, mockup generation, archive restoration.
  • → Open vision models (Qwen 2.5-VL, LLaVA, MoonDream): automatic image tagging, alt-text description, OCR.
  • → Redis queue + GPU workers: async jobs, retry policy, configurable priority (breaking news vs archive).
  • → Custom REST API facing the CMS: webhook on upload, status tracking, callback when ready.
  • → S3-compatible storage: separated input/output, configurable retention, no file ever leaves the EU perimeter.
Reference setup

Reference setup: 1 GPU server, ~30h of video transcribed per day.

1 dedicated GPU server (NVIDIA L40S or equivalent), Whisper large-v3 in batch, ComfyUI for images, overnight vision tagging. ~30 hours of video transcribed per day with diarisation, ~5k images tagged in the overnight batch. Fixed cost, no token-based pricing.

~30h
video/day
~5k
images/night
L40S
GPU
When we use it

Where we use it.

Free audit

Let’s talk about your stack,
free, no strings attached.

30 minutes with a Romiltec architect. Together we figure out whether Tech Performance is a fit, and if it isn’t, we tell you straight away. No cold pitch, no black-box quote.

Book a call cal.com/romiltec/tech-performance · 30 min call