Careers

Build the product behind liquid-cooled AI reliability.

Work on the reliability layer for liquid-cooled AI systems: coolant health, CDU telemetry, GPU thermal behavior, and the software that turns drift into action.

This is early, technical work for people who like physical systems, messy signals, and products that have to earn trust from operators.

Coolant chemistryTelemetryML systemsField reliability
Own real signals

Coolant, CDUs, racks, filters, flow, pressure, and GPU context.

Build the product

Baselines, diagnostics, APIs, alerts, and operator workflows.

Prove it works

Signals that help teams act before cooling drift hurts output.

People at workRole work zones
Reliability EngineWork happens where coolant, telemetry, and compute meet.

Each role owns a different signal, but the product only works when the signals explain one another.

Coolant sampleCDU telemetryGPU stateOperator action
Data scienceTelemetry to risk models

Coolant, CDU, rack, and GPU signals.

ChemistryFluid evidence

Particles, inhibitors, conductivity, and corrosion risk.

Field reliabilityReal loop checks

Sampling, inspection, service history, and recovery proof.

Software productOperator workflows

Alerts, APIs, timelines, and decision support.

Why it matters

The loop is now part of the compute stack.

Dense GPU systems depend on coolant health, flow, pressure, service history, and thermal margin. We are building the layer that reads those signals together.

Early signal

Catch drift before it becomes an incident

Read coolant, hydraulic, thermal, service, and workload signals as one operating picture.

Real systems

Turn science into product evidence

Make chemistry and field observations usable by models, alerts, reports, and customer decisions.

Useful action

Build for trust, not noise

Every recommendation should show evidence, confidence, and the next practical check.

What you will build

Build the reliability layer operators can trust.

Every role owns a real part of the product: what we measure, how we model it, how we explain it, and how a data center team acts on it.

Fluid

Make chemistry visible

Turn pH, conductivity, particles, turbidity, inhibitors, and contamination patterns into live operating signals.

Systems

Connect the loop

Bring CDU, rack, pressure, flow, filter, service, and GPU context into reliability views and APIs.

Models

Find quiet drift

Create baselines, anomaly logic, confidence windows, and explanations that survive noisy infrastructure data.

Workflow

Help operators move

Support inspection, sampling, rebalancing, conditioning, cleaning, output protection, and recovery checks.

Work surface

No single discipline owns the whole problem.

The useful answers come from chemistry, telemetry, thermal behavior, service history, and production software sitting in the same room.

Chemistry and materials

Fluid health, inhibitors, particles, contamination, corrosion risk, and materials compatibility.

Signal systems

Baselines, anomaly detection, feature engineering, telemetry quality, and reliability metrics.

Thermal infrastructure

CDUs, manifolds, filters, cold plates, pressure, flow, and cooling behavior.

Operator decisions

Clear next actions that help teams inspect, sample, rebalance, clean, or protect output earlier.

Role ownership

Each role owns a real part of the reliability product.

The work is not generic AI infrastructure. It is the chain from physical evidence to a decision an operator can trust.

Data Scientist
Own
Healthy baselines
Use
Python, SQL, time-series
Prove
Fewer false positives
Machine Learning Engineer - Predictive Reliability Systems
Own
Production reliability models
Use
Python or TypeScript pipelines
Prove
Explainable alerts
Coolant Chemistry & Materials Engineer
Own
Coolant evidence
Use
Chemistry and materials
Prove
Field-ready limits
Field Reliability Engineer - Liquid Cooling Systems
Own
Operator workflows
Use
Telemetry and physical checks
Prove
Recovery after action

Open role areas

Small team. Serious systems.

We hire around ownership. A strong fit might come from lab work, data systems, controls, software, thermal systems, or field reliability.

Remote / San Francisco, CA

Data Scientist

FULL TIME

Own the analysis and models that make loop behavior readable: define normal, catch drift, and explain it in terms an operator can trust.

Best fit if you have worked with real telemetry, anomaly detection, and model explanations where false positives matter.

Own

Healthy baselines

Use

Python, SQL, time-series

Prove

Fewer false positives

View role details

ImpactYour work helps teams know when to inspect, sample, rebalance, clean, or protect workload output before cooling margin disappears.

First problems
  • Define healthy baseline windows for coolant condition, pressure, flow, delta T, service events, and GPU thermal response.
  • Build drift views that separate workload movement from loop restriction, fluid degradation, maintenance events, and sensor noise.
  • Design labels, validation plans, and feedback loops with chemistry, field reliability, and product teams.
Skills we look for
  • Strong Python and SQL for cleaning data, building features, and making analysis reproducible.
  • Experience with time-series data, anomaly detection, statistical baselines, model validation, or risk scoring.
  • Comfort with missing values, calibration drift, outliers, service resets, and changing operating regimes.
  • Ability to explain uncertainty plainly to software, field, chemistry, thermal, and customer-facing teams.
Strong signals
  • You have shipped analysis that changed an operational decision, not only a dashboard.
  • You care about precision, recall, and the cost of a wrong alert.
  • You enjoy physical systems and can work without perfect labels on day one.

Remote / San Francisco, CA

Machine Learning Engineer - Predictive Reliability Systems

FULL TIME

Build the production ML systems behind coolant and loop reliability: ingestion, features, models, explanations, monitoring, and APIs.

Best fit if you can move between model design, production software, observability, and practical reliability tradeoffs.

Own

Production reliability models

Use

Python or TypeScript pipelines

Prove

Explainable alerts

View role details

ImpactThe product should not behave like a black-box score. It should say what changed, how confident it is, what evidence supports it, and what action is worth taking.

First problems
  • Design pipelines that combine coolant health, CDU behavior, thermal response, maintenance context, and workload state.
  • Build explainability, model monitoring, data-quality checks, and versioned evaluation for field validation.
  • Turn predictions into APIs and operator workflows without creating alert fatigue.
Skills we look for
  • Strong Python and/or TypeScript, with experience shipping ML or data systems beyond notebooks.
  • Practical judgment across batch or streaming pipelines, model evaluation, model monitoring, and data versioning.
  • Experience with time-series forecasting, anomaly detection, classification, ranking, or probabilistic risk scoring.
  • Good instincts around precision, recall, alert thresholds, explainability, and when a model should defer instead of guessing.
Strong signals
  • You write production code and still care deeply about model behavior.
  • You can debug a pipeline, a bad label, and a confusing operator experience in the same week.
  • You have seen industrial telemetry, observability, controls, digital twins, reliability, or infrastructure data.

Remote / San Francisco, CA

Coolant Chemistry & Materials Engineer

FULL TIME

Turn coolant chemistry, materials compatibility, contamination, and degradation into reliability evidence that infrastructure teams can use.

Best fit if you can move from fluid science to practical limits, sampling plans, and product-facing interpretation.

Own

Coolant evidence

Use

Chemistry and materials

Prove

Field-ready limits

View role details

ImpactYou will connect lab evidence, side-stream sensing, field samples, and operating context so coolant health becomes an operational signal, not a disconnected report.

First problems
  • Define which coolant measurements matter, how often they matter, and what movement should change an operator decision.
  • Map chemistry changes to risks such as corrosion, fouling, deposits, filter loading, biological growth, and heat-transfer loss.
  • Create validation plans that help data teams build labels, thresholds, and confidence around coolant-health models.
Skills we look for
  • Background in chemistry, materials science, chemical engineering, corrosion, coolant formulation, water treatment, or fluid reliability.
  • Working knowledge of pH, conductivity, turbidity, particles, inhibitors, organic acids, microbial risk, contamination, and corrosion mechanisms.
  • Ability to design test plans, sampling protocols, acceptance windows, and failure-analysis workflows.
  • Comfort collaborating with software and data teams so chemistry becomes structured data, not only lab notes.
Strong signals
  • You can explain chemistry in a way operators and software teams can act on.
  • You know where lab certainty ends and field judgment begins.
  • You have worked with glycol/water loops, CDUs, cold plates, data centers, semiconductors, or industrial cooling.

Remote / San Francisco, CA

Field Reliability Engineer - Liquid Cooling Systems

FULL TIME

Bring real field behavior into the product so Reliability Engine reflects how liquid-cooled systems are installed, operated, serviced, and recovered.

Best fit if you can read telemetry, reason from the physical system, and write procedures that people will actually follow.

Own

Operator workflows

Use

Telemetry and physical checks

Prove

Recovery after action

View role details

ImpactYou will help make recommendations technically correct and usable during real decisions: diagnose drift, check the right hardware, and verify recovery after action.

First problems
  • Turn inspection, sampling, balancing, filter changes, cleaning, and recovery checks into product workflows.
  • Review abnormal behavior around maintenance events, pump changes, filter loading, coolant conditioning, and thermal response.
  • Validate whether alerts and recommendations match what field teams can safely check or do next.
Skills we look for
  • Experience with data center infrastructure, liquid cooling, mechanical systems, controls, thermal systems, reliability, or field engineering.
  • Ability to read telemetry and connect it to physical causes: restriction, imbalance, pump behavior, filter loading, air, fouling, or sensor problems.
  • Clear writing for procedures, root-cause notes, customer-facing explanations, and engineering handoffs.
  • Comfort working across hardware, software, operations, and customer environments where data is incomplete.
Strong signals
  • You have commissioned, supported, troubleshot, or operated systems where uptime mattered.
  • You know the difference between an elegant recommendation and one a technician can execute.
  • You have seen direct-to-chip cooling, CDU commissioning, GPU clusters, facilities operations, or incident/postmortem work.

Before you reach out

Show the real systems behind your work.

A good note is short: what you built, what signals you handled, and which reliability problem you want to work on next.

Are these roles open now?

These are active role areas. We are especially interested in people with chemistry, thermal systems, telemetry, ML, controls, sensors, reliability engineering, or data center infrastructure experience.

What experience is most relevant?

Relevant work includes time-series data, anomaly detection, fluid testing, thermal systems, field reliability, controls, high-density infrastructure, and products used by operators.

What should I include when applying?

Use the Apply for this role button and share the systems you worked on, the signals or tools you handled, and the reliability problem you want to solve next.