Coolant, CDUs, racks, filters, flow, pressure, and GPU context.
Careers
Build the product behind liquid-cooled AI reliability.
Work on the reliability layer for liquid-cooled AI systems: coolant health, CDU telemetry, GPU thermal behavior, and the software that turns drift into action.
This is early, technical work for people who like physical systems, messy signals, and products that have to earn trust from operators.
Baselines, diagnostics, APIs, alerts, and operator workflows.
Signals that help teams act before cooling drift hurts output.
Why it matters
The loop is now part of the compute stack.
Dense GPU systems depend on coolant health, flow, pressure, service history, and thermal margin. We are building the layer that reads those signals together.
Catch drift before it becomes an incident
Read coolant, hydraulic, thermal, service, and workload signals as one operating picture.
Turn science into product evidence
Make chemistry and field observations usable by models, alerts, reports, and customer decisions.
Build for trust, not noise
Every recommendation should show evidence, confidence, and the next practical check.
What you will build
Build the reliability layer operators can trust.
Every role owns a real part of the product: what we measure, how we model it, how we explain it, and how a data center team acts on it.
Make chemistry visible
Turn pH, conductivity, particles, turbidity, inhibitors, and contamination patterns into live operating signals.
Connect the loop
Bring CDU, rack, pressure, flow, filter, service, and GPU context into reliability views and APIs.
Find quiet drift
Create baselines, anomaly logic, confidence windows, and explanations that survive noisy infrastructure data.
Help operators move
Support inspection, sampling, rebalancing, conditioning, cleaning, output protection, and recovery checks.
Work surface
No single discipline owns the whole problem.
The useful answers come from chemistry, telemetry, thermal behavior, service history, and production software sitting in the same room.
Chemistry and materials
Fluid health, inhibitors, particles, contamination, corrosion risk, and materials compatibility.
Signal systems
Baselines, anomaly detection, feature engineering, telemetry quality, and reliability metrics.
Thermal infrastructure
CDUs, manifolds, filters, cold plates, pressure, flow, and cooling behavior.
Operator decisions
Clear next actions that help teams inspect, sample, rebalance, clean, or protect output earlier.
Role ownership
Each role owns a real part of the reliability product.
The work is not generic AI infrastructure. It is the chain from physical evidence to a decision an operator can trust.
- Own
- Healthy baselines
- Use
- Python, SQL, time-series
- Prove
- Fewer false positives
- Own
- Production reliability models
- Use
- Python or TypeScript pipelines
- Prove
- Explainable alerts
- Own
- Coolant evidence
- Use
- Chemistry and materials
- Prove
- Field-ready limits
- Own
- Operator workflows
- Use
- Telemetry and physical checks
- Prove
- Recovery after action
Open role areas
Small team. Serious systems.
We hire around ownership. A strong fit might come from lab work, data systems, controls, software, thermal systems, or field reliability.
Remote / San Francisco, CA
Data Scientist
Own the analysis and models that make loop behavior readable: define normal, catch drift, and explain it in terms an operator can trust.
Best fit if you have worked with real telemetry, anomaly detection, and model explanations where false positives matter.
Healthy baselines
Python, SQL, time-series
Fewer false positives
View role details
ImpactYour work helps teams know when to inspect, sample, rebalance, clean, or protect workload output before cooling margin disappears.
- Define healthy baseline windows for coolant condition, pressure, flow, delta T, service events, and GPU thermal response.
- Build drift views that separate workload movement from loop restriction, fluid degradation, maintenance events, and sensor noise.
- Design labels, validation plans, and feedback loops with chemistry, field reliability, and product teams.
- Strong Python and SQL for cleaning data, building features, and making analysis reproducible.
- Experience with time-series data, anomaly detection, statistical baselines, model validation, or risk scoring.
- Comfort with missing values, calibration drift, outliers, service resets, and changing operating regimes.
- Ability to explain uncertainty plainly to software, field, chemistry, thermal, and customer-facing teams.
- You have shipped analysis that changed an operational decision, not only a dashboard.
- You care about precision, recall, and the cost of a wrong alert.
- You enjoy physical systems and can work without perfect labels on day one.
Remote / San Francisco, CA
Machine Learning Engineer - Predictive Reliability Systems
Build the production ML systems behind coolant and loop reliability: ingestion, features, models, explanations, monitoring, and APIs.
Best fit if you can move between model design, production software, observability, and practical reliability tradeoffs.
Production reliability models
Python or TypeScript pipelines
Explainable alerts
View role details
ImpactThe product should not behave like a black-box score. It should say what changed, how confident it is, what evidence supports it, and what action is worth taking.
- Design pipelines that combine coolant health, CDU behavior, thermal response, maintenance context, and workload state.
- Build explainability, model monitoring, data-quality checks, and versioned evaluation for field validation.
- Turn predictions into APIs and operator workflows without creating alert fatigue.
- Strong Python and/or TypeScript, with experience shipping ML or data systems beyond notebooks.
- Practical judgment across batch or streaming pipelines, model evaluation, model monitoring, and data versioning.
- Experience with time-series forecasting, anomaly detection, classification, ranking, or probabilistic risk scoring.
- Good instincts around precision, recall, alert thresholds, explainability, and when a model should defer instead of guessing.
- You write production code and still care deeply about model behavior.
- You can debug a pipeline, a bad label, and a confusing operator experience in the same week.
- You have seen industrial telemetry, observability, controls, digital twins, reliability, or infrastructure data.
Remote / San Francisco, CA
Coolant Chemistry & Materials Engineer
Turn coolant chemistry, materials compatibility, contamination, and degradation into reliability evidence that infrastructure teams can use.
Best fit if you can move from fluid science to practical limits, sampling plans, and product-facing interpretation.
Coolant evidence
Chemistry and materials
Field-ready limits
View role details
ImpactYou will connect lab evidence, side-stream sensing, field samples, and operating context so coolant health becomes an operational signal, not a disconnected report.
- Define which coolant measurements matter, how often they matter, and what movement should change an operator decision.
- Map chemistry changes to risks such as corrosion, fouling, deposits, filter loading, biological growth, and heat-transfer loss.
- Create validation plans that help data teams build labels, thresholds, and confidence around coolant-health models.
- Background in chemistry, materials science, chemical engineering, corrosion, coolant formulation, water treatment, or fluid reliability.
- Working knowledge of pH, conductivity, turbidity, particles, inhibitors, organic acids, microbial risk, contamination, and corrosion mechanisms.
- Ability to design test plans, sampling protocols, acceptance windows, and failure-analysis workflows.
- Comfort collaborating with software and data teams so chemistry becomes structured data, not only lab notes.
- You can explain chemistry in a way operators and software teams can act on.
- You know where lab certainty ends and field judgment begins.
- You have worked with glycol/water loops, CDUs, cold plates, data centers, semiconductors, or industrial cooling.
Remote / San Francisco, CA
Field Reliability Engineer - Liquid Cooling Systems
Bring real field behavior into the product so Reliability Engine reflects how liquid-cooled systems are installed, operated, serviced, and recovered.
Best fit if you can read telemetry, reason from the physical system, and write procedures that people will actually follow.
Operator workflows
Telemetry and physical checks
Recovery after action
View role details
ImpactYou will help make recommendations technically correct and usable during real decisions: diagnose drift, check the right hardware, and verify recovery after action.
- Turn inspection, sampling, balancing, filter changes, cleaning, and recovery checks into product workflows.
- Review abnormal behavior around maintenance events, pump changes, filter loading, coolant conditioning, and thermal response.
- Validate whether alerts and recommendations match what field teams can safely check or do next.
- Experience with data center infrastructure, liquid cooling, mechanical systems, controls, thermal systems, reliability, or field engineering.
- Ability to read telemetry and connect it to physical causes: restriction, imbalance, pump behavior, filter loading, air, fouling, or sensor problems.
- Clear writing for procedures, root-cause notes, customer-facing explanations, and engineering handoffs.
- Comfort working across hardware, software, operations, and customer environments where data is incomplete.
- You have commissioned, supported, troubleshot, or operated systems where uptime mattered.
- You know the difference between an elegant recommendation and one a technician can execute.
- You have seen direct-to-chip cooling, CDU commissioning, GPU clusters, facilities operations, or incident/postmortem work.
Before you reach out
Show the real systems behind your work.
A good note is short: what you built, what signals you handled, and which reliability problem you want to work on next.
Are these roles open now?
These are active role areas. We are especially interested in people with chemistry, thermal systems, telemetry, ML, controls, sensors, reliability engineering, or data center infrastructure experience.
What experience is most relevant?
Relevant work includes time-series data, anomaly detection, fluid testing, thermal systems, field reliability, controls, high-density infrastructure, and products used by operators.
What should I include when applying?
Use the Apply for this role button and share the systems you worked on, the signals or tools you handled, and the reliability problem you want to solve next.