--- title: 'Senior Modeling Architect, Performance Benchmarking at Neurophos' canonical: 'https://feeny.ai/job/senior-modeling-architect-performance-benchmarking-neurophos-austin-rpy828cj7yrc' type: 'job' last_seen: '2026-09-07' --- # Senior Modeling Architect, Performance Benchmarking at Neurophos - **Company:** Neurophos - **Location:** Austin, TX - **Compensation:** $210k–$250k - **Employment:** full-time - **Work type:** onsite - **Posted:** 2026-09-04 - **Last confirmed live:** 2026-09-07 - **Apply:** https://jobs.ashbyhq.com/neurophos/e97c4a7a-c36e-4d29-8c8c-2baad33cd04c ## Job description ## ABOUT NEUROPHOS The demand for new data centers and AI compute is rapidly outpacing the planet's energy capacity. Digital solutions are hitting a power wall as we approach the physical limits of traditional silicon. Conquering this bottleneck means rethinking the fundamental architecture of inference compute. The industry's current path can't meet the need, so we're taking a different approach. Instead of traditional electronic circuits, we use silicon photonics and an active, programmable metasurface to perform matrix multiplications at the speed of light. Our optical cells are 10,000x smaller than traditional photonic components, enabling unprecedented density. By using photonics instead of electricity, our chips become more efficient as they scale. This architecture will deliver up to 100 times the energy efficiency of existing solutions while significantly improving performance for large-scale AI inference. We’ve assembled a world-class team of industry veterans and recently raised a $110M Series A https://www.neurophos.com/110m-raise led by Gates Frontier. Participants include M12 (Microsoft’s Venture Fund), Carbon Direct Capital, Aramco Ventures, Bosch Ventures, Tectonic Ventures, Space Capital, and others. Join us and shape the future of computing! Location: Austin, TX or Sunnyvale, CA. Full-time onsite position. Reports To: Sr. Director of Modeling FLSA Status: Exempt ## POSITION OVERVIEW We are seeking a performance engineer to own the benchmarking numbers behind the T100 optical inference accelerator. Architecture and product decisions here are made on measured performance and energy, and this role produces those figures for the same workloads at every level of fidelity we use: roofline and limiter analysis, architecture performance models, in-house RTL simulation, and measured runs on competing GPUs and accelerators. You will join the Architecture and Modeling team, set the measurement methodology, and keep it current as models, software stacks, drivers, and hardware generations turn over. Every result ships with the harness, config, plots, logs, and assumptions behind it, so anyone can rerun it and see how the number was reached. ## KEY RESPONSIBILITIES - Own the performance and energy metrics that architecture, product, and leadership rely on, and keep them consistent across modeling fidelities and measured hardware. - Produce numbers for the same workloads across all fidelities we use internally: roofline and limiter models, architecture performance models, in-house RTL simulation, and measured competitor hardware. Work with the modeling team to keep the simulated and measured workload sets aligned. - Hold workload definitions constant across fidelities, including model or application, sequence length, batch, precision, prefill versus decode, and tensor, pipeline, and sequence parallelism. - Bring up inference workloads from Hugging Face, PyTorch, published papers, and vendor stacks such as vLLM, SGLang, TensorRT-LLM, and Triton Inference Server. These include dense and Mixture-of-Experts (MoE) transformers, attention and KV-cache reduction strategies, and hybrid/SSM models, as well as retrieval, speech, vision, and recommendation workloads that map onto the accelerator. - Measure competing GPUs and accelerators end-to-end, owning the cloud or lab account, image, drivers, and run recipe. - Report time to first token (TTFT), inter-token latency (ITL), tokens per second, tokens per second per watt, and energy, using nvidia-smi, DCGM, power capping, or equivalent instrumentation. - Document where RTL simulation, the performance model, and measured competitor results disagree, and attach the configs and logs behind each. - Maintain a reviewed internal benchmark suite. Keep internal-only results clearly separate from anything cleared for customer or public use, and route external claims through the designated approver before they ship. ## QUALIFICATIONS - BS or MS in Computer Engineering, Electrical Engineering, Computer Science, or equivalent practical experience. - 5+ years of experience in GPU performance engineering, accelerator benchmarking, HPC performance measurement, or ML systems measurement. - Track record of building or operating benchmark harnesses that produced measured results on real GPUs or accelerators, including turning a Hugging Face model card, paper, or application description into a runnable benchmark. - Hands-on experience with roofline analysis, limiter analysis, or analytical performance modeling. - GPU performance analysis with NVIDIA Nsight Systems and Nsight Compute, or an equivalent profiler, covering HBM-bound versus compute-bound analysis, precision (FP16, BF16, FP8, INT8), and batching. - Working knowledge of LLM inference stacks such as Hugging Face, vLLM, SGLang, or TensorRT-LLM, including prefill versus decode, continuous batching, and MoE. - Proficiency in Python for harnesses, parsing, and plots, and comfort working in Linux. - Cloud GPU operations on AWS, GCP, or Azure, including containers, instance types, drivers, quotas, and cost. ## PREFERRED SKILLS - Experience correlating a performance model or RTL/Verilator simulation against measured silicon or GPUs. - GPU kernel work in CUDA, CUTLASS, or Triton, or familiarity with PyTorch internals. - Familiarity with current inference-serving internals such as PagedAttention, FlashAttention, speculative decoding, and disaggregated prefill. - Distributed inference experience covering collectives, all-reduce, NCCL, NVLink, and InfiniBand, or work with MLPerf or production inference benchmarking pipelines. - Background at a hyperscaler, GPU vendor, accelerator company, or inference lab. ## WHAT WE OFFER This is an opportunity to play a pivotal role in an innovative startup redefining the future of AI hardware. Work on game-changing technology at the intersection of photonics and AI as part of a collaborative, brilliant team. You’ll contribute to a platform that redefines computational performance and accelerates the future of artificial intelligence. Come help us bring this transformative technology to the world. ## BENEFITS Join a team that invests in your future and your well-being. At Neurophos, we offer: - 100% coverage of base health plan premiums for you and your dependents, plus HSA contributions. - Unlimited PTO. No rigid vacation banks, just a focus on delivery. - 401(k) matching and stock option opportunities to ensure our success is your success. - Full suite of voluntary benefits, including Dental, Vision, Life, Hospital, Critical Illness, and Accident insurance. - Personalized Benefits. Choose the plans that fit your life and take the cash back for those that don’t. ## About Neurophos ## Company Overview - **One-liner**: Neurophos develops photonic AI chips (Optical Processing Units) that perform matrix multiplications at the speed of light, aiming to replace traditional GPUs in data-center inference workloads. - **Entity Type**: Private, Series A - **Headquarters**: Austin, Texas, USA (with an additional office in San Jose, California) - **Founded**: 2020 - **Founders**: Patrick Bowen (CEO, Co-Founder) and Andrew Traverso (Chief Scientist, Co-Founder); other co-founders may include Preston Woo and Hod Finkelstein (CEO, CTO roles indicated on team page). ## Core Business - **Primary industry**: Semiconductor / Photonic AI Hardware - **Target customers**: B2B – hyperscale data centers, enterprise AI infrastructure providers - **Mission/purpose**: “We’re building the first programmable light-speed computer” – solving energy and scalability challenges in AI data centers by using optical rather than electronic computation. ## Products & Services - **Optical Processing Unit (OPU)**: A photonic AI chip that integrates over one million micron-scale optical processing components on a single die. It performs matrix multiplications in-memory at the speed of light, delivering up to 100x the energy efficiency of traditional GPUs and NVIDIA servers while fitting in a single-GPU form factor. The chip uses an active, programmable metasurface and silicon photonics to achieve 0.47 ExaOPS performance in a 1m x 1m footprint. First systems are targeted for early 2028, with production ramp in mid-2028. ## Market Standing - **Valuation/Market Cap**: Not disclosed - **Key Metric**: Total funding raised = **$119.2M** (including a $110M Series A closed in early 2026) - **Notable Investors**: Gates Frontier (lead), M12 (Microsoft’s venture fund), Carbon Direct Capital, Aramco Ventures, Bosch Ventures, Tectonic Ventures, Space Capital, MetaVC Partners, Gaingels, Mana Ventures, and 22+ others. - **Growth Signals**: - Recognized on EE Times Silicon 100 list for several consecutive years. - Over 300 patents filed and protected. - Rapid hiring after Series A – “hypergrowth phase” as per press release. - Team includes veterans from NVIDIA, AMD, Apple, Meta, Google, Magic Leap, Micron, Qualcomm, Samsung, and other photonic/AR companies. ## Competitive Advantages - **Metasurface optical transistors** that are **10,000x smaller** than traditional photonic components, enabling unprecedented density – millions of weights fit in the area of a postage stamp. - **300+ patents** covering optical metamaterials, neural network accelerators, and machine learning hardware. - **Energy efficiency**: Up to 100x improvement over existing electronic solutions (e.g., GPUs), directly addressing the power constraints of hyperscale AI inference. - **Size/performance**: OPU delivers the performance of a 3,000-pound server in the size and power envelope of a single GPU. ## Strategic Focus - **Commercialization**: First customer evaluations begin in 2026; first OPU systems target early 2028 production ramp. - **Talent acquisition**: Scaling the team aggressively across analog design, photonics, architecture, and software to meet “hypergrowth” demand. - **Partnerships**: Leveraging investors like Bosch Ventures (industrial IoT), Aramco Ventures (energy), and M12 (cloud/AI) to align with large-scale data center customers. ## Why Work Here - **Culture**: In-office, on‑site work culture in Austin, TX (primary) and San Jose, CA. Described as a “world-class team” with leaders from top semiconductor and tech companies. - **Growth**: Joining at a hypergrowth scale‑up stage with significant funding and a clear path to production; opportunities to shape the architecture and engineering of a new computing paradigm. - **Engineering focus**: 34 out of 40 total employees are in product + tech roles, indicating a highly technical, R&D-driven environment. - **Impact**: Work on cutting-edge optical computing that could redefine AI hardware efficiency; involvement in everything from silicon photonics to compiler software. - **Notable perks**: Not explicitly listed, but being a well-funded deep-tech startup likely offers competitive equity, technical challenges, and direct mentorship from industry veterans. ## Sources 1. [neurophos.com](https://www.neurophos.com/) 2. [neurophos.com/careers](https://www.neurophos.com/careers) 3. [neurophos.com/team](https://www.neurophos.com/team) 4. [cbinsights.com](https://www.cbinsights.com/company/neurophos) 5. [builtin.com](https://builtin.com/company/neurophos) ## Other roles at Neurophos - [Staff Modeling Architect](https://feeny.ai/job/staff-modeling-architect-neurophos-austin-8cedvv6953bg) — Austin, TX - [Modeling Architect](https://feeny.ai/job/modeling-architect-neurophos-austin-ew2q0rxv7e3b) — Austin, TX - [Principal SoC Design Engineer / SoC Lead](https://feeny.ai/job/principal-soc-design-engineer-soc-lead-neurophos-austin-ghf226b0sp4d) — Austin, TX - [Physical Design / Implementation Lead](https://feeny.ai/job/physical-design-implementation-lead-neurophos-austin-jge0dfb6md1v) — Austin, TX - [Senior/Staff Applied Scientist, Numerical Optimization & Quantization](https://feeny.ai/job/senior-staff-applied-scientist-numerical-optimization-quantization-neurophos-jgmrpeawfsp5) — Austin, TX - [Principal Fabric/Network Architect](https://feeny.ai/job/principal-fabric-network-architect-neurophos-sunnyvale-zq97mx2rm8r3) — Sunnyvale, CA - [Principal Platform Security Architect](https://feeny.ai/job/principal-platform-security-architect-neurophos-sunnyvale-fw052b83t93h) — Sunnyvale, CA - [Lead IT Engineer](https://feeny.ai/job/lead-it-engineer-neurophos-austin-jegb2gkjwpgy) — Austin, TX - [Lead Security Engineer](https://feeny.ai/job/lead-security-engineer-neurophos-austin-er5qq4nwr4ze) — Austin, TX - [Director of Facilities](https://feeny.ai/job/director-of-facilities-neurophos-austin-jey9dppssncz) — Austin, TX