Whisper for audio/video transcription, ComfyUI for image workflows, vision models (Qwen 2.5-VL, LLaVA) for automatic tagging. All on dedicated EU GPUs, overnight batch integrated into the CMS, no uploads to third-party cloud vendors.
1 dedicated GPU server (NVIDIA L40S or equivalent), Whisper large-v3 in batch, ComfyUI for images, overnight vision tagging. ~30 hours of video transcribed per day with diarisation, ~5k images tagged in the overnight batch. Fixed cost, no token-based pricing.
Italian audio transcribed and diarised for the journalist, article turnaround cut in half.
SRT/VTT generation for editorial video, timeline alignment, optional multi-language translation.
Thousands of historical images, subject/place recognition, semantic search for the newsroom.
Alt-text generation for WCAG accessibility, on the already-published archive.
Thumbnail variants in every format (16:9, 1:1, 9:16) for multi-platform CMS and social.
NSFW/violence detection on user-generated content before editorial go-live.
30 minutes with a Romiltec architect. Together we figure out whether Tech Performance is a fit, and if it isn’t, we tell you straight away. No cold pitch, no black-box quote.