--- title: 'Vision Language Model Engineer at EchoTwin AI' canonical: 'https://feeny.ai/job/vision-language-model-engineer-echotwin-ai-san-francisco-fkagxqaz332s' type: 'job' last_seen: '2026-09-11' --- # Vision Language Model Engineer at EchoTwin AI - **Company:** EchoTwin AI - **Location:** San Francisco, CA - **Employment:** full-time - **Work type:** onsite - **Posted:** 2025-09-06 - **Last confirmed live:** 2026-09-11 - **Apply:** https://jobs.ashbyhq.com/echotwin/77fae377-e76d-4146-ac55-06a84af983ef ## Job description Company Overview EchoTwin AI is pioneering AI-driven infrastructure intelligence, redefining how cities are managed. Powered by a proprietary visual intelligence engine with full spatial reasoning, EchoTwin transforms municipal fleets into mobile urban sensors—creating living digital twins that provide real-time insights into infrastructure, compliance, and safety. By enabling municipalities to proactively monitor, predict, and resolve issues, EchoTwin helps build resilient, self-healing, and sustainable urban ecosystems. More than “smart cities,” EchoTwin is advancing the era of cognizant cities—urban environments with the awareness to see, think, and act on challenges in real time. ## What You’ll Do As a Vision Language Model Engineer, you will design, develop, and optimize advanced vision-language models that integrate visual and textual data to enable intelligent systems. You will work closely with cross-functional teams to build models that power applications such as image captioning, visual question answering, and multimodal AI at the edge. ## Key Responsibilities - Design and implement state-of-the-art vision-language models using deep learning frameworks. - Develop and fine-tune models that combine computer vision and natural language processing for tasks like image captioning, visual question answering, and text-to-image generation. - Collaborate with data scientists and software engineers to integrate models into production systems. - Optimize model performance for accuracy, latency, and scalability in real-world applications. - Conduct experiments to evaluate model performance and iterate on architectures and training pipelines. - Stay up-to-date with the latest research in vision-language models and incorporate advancements into projects. - Contribute to data preprocessing, augmentation, and annotation pipelines for multimodal datasets. - Document model development processes and present findings to technical and non-technical stakeholders. ## Qualifications - Bachelor’s, Master’s or Ph.D. in Computer Science, Machine Learning, Artificial Intelligence, or a related field (or equivalent experience). - 3+ years of experience in machine learning, with a focus on vision-language models or multimodal AI. - Hands-on experience with deep learning frameworks such as PyTorch or TensorFlow. - Proven track record of building and deploying computer vision and/or NLP models. - Proficiency in Python and relevant ML libraries (e.g., Hugging Face, OpenCV, Transformers). - Experience with large-scale model training and optimization (e.g., distributed training, quantization). - Strong understanding of neural network architectures (e.g., CNNs, Transformers, CLIP, or similar). - Experience with multimodal datasets and preprocessing techniques for images and text. - Familiarity with cloud platforms (e.g., AWS, GCP, Azure) and model deployment workflows. - Strong problem-solving skills and ability to work in a fast-paced, collaborative environment. - Excellent communication skills to explain complex technical concepts to diverse audiences. ## Benefits and Perks There are endless learning and development opportunities from a highly diverse and talented peer group, including experts in various fields, including Computer Vision, GenAI, Digital Twin, Government Contracting, Systems and Device Engineering, Operations, Communications, and more! - Options for medical, dental, and vision coverage for employees and dependents (for US employees) - Flexible Spending Account (FSA) and Dependent Care Flexible Spending Account (DCFSA) - 401(k) with 3% company matching - Unlimited PTO - Profit sharing Please do not forward resumes to our jobs alias, EchoTwin AI employees, or any other company location. EchoTwin AI is not responsible for any fees related to unsolicited resumes. Life at EchoTwin AI If you want to empower the world’s most important cities—and the institutions that run them—you belong here. At EchoTwin AI, we value excellence regardless of background and are committed to building a team that reflects the communities we serve. EchoTwin AI is an Equal Opportunity Employer. We consider all qualified applicants without regard to race, color, religion, sex (including pregnancy), sexual orientation, gender identity or expression, national origin, age, disability, veteran status, genetic information, or any other status protected by applicable law. ## About EchoTwin AI ## Company Overview - **One-liner**: EchoTwin AI builds a Physical AI operating system for cities, turning municipal fleets and infrastructure into intelligent, self-aware urban networks. - **Entity Type**: Private (Seed stage, $25M total funding) - **Headquarters**: Boca Raton, Florida, United States (also offices in San Francisco, New York, Dubai, Abu Dhabi, Riyadh, Doha) - **Founded**: 2024 - **Founders**: Chris Carson (Founder & CEO) ## Core Business - **Primary industry**: Urban AI / Smart City Infrastructure / Computer Vision - **Target customers**: Municipal governments, city authorities, urban planning departments (B2G/B2B) - **Mission**: To enhance urban safety, sustainability, and resilience through AI and geospatial intelligence, enabling self-healing cities that can see, understand, and respond in real time. ## Products & Services - **[CityView](https://www.echotwin.ai/)**: A vision-language model (VLM) purpose-built for cities. It interprets complex visual and textual data to detect infrastructure issues, compliance violations, and hazards in real time. Deployed via edge devices on vehicles, drones, and autonomous systems. - **[EchoTwin Platform](https://www.echotwin.ai/)**: A vertically integrated edge-to-cloud platform that ingests data from municipal sensors, runs AI analytics (predictive, diagnostic, geospatial), and provides dashboards, APIs, and automated workflows for city operators. ## Market Standing - **Valuation/Market Cap**: Not disclosed (private) - **Key Metric**: $25M total funding (two seed rounds: $17M led by Metis Ventures in Feb 2025, $8M in Sep 2025) - **Notable Investors/Partners**: Metis Ventures (lead), plus 6 other investors per Seed round; partnerships with pioneering cities (e.g., NEOM, Dubai 2040, Paris 2050 initiatives) - **Growth Signals**: 106.7% headcount growth YoY (28 employees); 2,368 LinkedIn followers (+223% yearly); active job postings (9 open roles); presence in 5 countries (US, Serbia, UAE, Greece, Portugal) ## Competitive Advantages - **Proprietary CityView VLM**: A specialized vision-language model trained on urban contexts, providing human-level perception for infrastructure monitoring. - **Edge-to-cloud architecture**: Real-time processing on vehicles/drones reduces latency and bandwidth needs, enabling autonomous corrective actions. - **Regulatory compliance & ethics**: Adheres to NIST AI Risk Management Framework, ETSI EN 303 645 cybersecurity standard, GDPR, and ISO/IEC 27701 – a key trust differentiator for government contracts. - **Self-healing city concept**: Moves beyond detection to automated remediation via agentic workflows, offering a proactive vs. reactive urban management model. ## Strategic Focus - **Global expansion**: Targeting smart city initiatives in North America, Middle East, Europe, and Asia (NEOM, Dubai 2040, Paris 2050). - **Product development**: Advancing CityView capabilities (geospatial reasoning, anomaly detection, predictive analytics) and expanding API/SDK ecosystem. - **Scaling deployments**: Converting municipal fleets into mobile sensing networks and integrating with third-party urban systems. ## Why Work Here - **Mission-driven impact**: Work on tangible urban challenges (safety, sustainability, resilience) with real-world outcomes. - **Global, distributed team**: Offices in San Francisco, New York, Dubai, and remote-friendly roles; culture of collaboration and mutual empowerment. - **Engineering focus**: 35% of staff in technical roles (CV, VLM, firmware, data science); cutting-edge stack (computer vision, LLMs, edge computing). - **Growth trajectory**: Early-stage (founded 2024) with rapid headcount growth and $25M funding – opportunity to shape product and culture. - **Perks & culture**: Emphasis on responsible AI, ethical practices, and a tight-knit, nimble team. Some roles are in-office (San Francisco, New York, Dubai) while others offer remote flexibility. ## Sources 1. [EchoTwin AI Website](https://www.echotwin.ai/) 2. [EchoTwin AI About Page](https://www.echotwin.ai/about) 3. [EchoTwin AI Careers Page](https://www.echotwin.ai/careers) 4. [EchoTwin AI LinkedIn](https://www.linkedin.com/company/echotwinai) 5. [EchoTwin AI on Built In](https://builtin.com/company/echotwin-ai) 6. [EchoTwin AI Job Board (Ashby)](https://jobs.ashbyhq.com/echotwin) ## Other roles at EchoTwin AI - [Prognostics Engineer](https://feeny.ai/job/prognostics-engineer-echotwin-ai-san-francisco-0e8wjgp967am) — San Francisco, CA - [3D Perception Engineer - SLAM with Mono Cameras](https://feeny.ai/job/3d-perception-engineer-slam-with-mono-cameras-echotwin-ai-san-francisco-n76t1fcdm16w) — San Francisco, CA - [Application Engineer – Computer Vision](https://feeny.ai/job/application-engineer-computer-vision-echotwin-ai-san-francisco-ynzgx65r8s3e) — San Francisco, CA - [Senior Data Scientist](https://feeny.ai/job/senior-data-scientist-echotwin-ai-san-francisco-zvshw5yzs5gr) — San Francisco, CA - [Machine Learning Engineer](https://feeny.ai/job/machine-learning-engineer-echotwin-ai-belgrade-vgc9tqcgzb4n) — Belgrade, Serbia - [Client Implementation Project Manager](https://feeny.ai/job/client-implementation-project-manager-echotwin-ai-dubai-ffx4rjwvcdy2) — Dubai, United Arab Emirates - [Software Engineer, Backend Developer](https://feeny.ai/job/software-engineer-backend-developer-echotwin-ai-belgrade-5yh4pwwmcdsr) — Belgrade, Serbia - [Senior Security Engineer](https://feeny.ai/job/senior-security-engineer-echotwin-ai-san-francisco-2tvmanh1s4zs) — San Francisco, CA - [Senior Technical Program Manager](https://feeny.ai/job/senior-technical-program-manager-echotwin-ai-san-francisco-qbtv703bm0de) — San Francisco, CA - [Senior Product Manager](https://feeny.ai/job/senior-product-manager-echotwin-ai-san-francisco-6ssr7eyvt5yj) — San Francisco, CA