Senior Distributed Systems / Platform Engineer
Software Engineering
San Francisco, CA, USA
You will be responsible for building the distributed infrastructure that allows Satlyt to securely operate hundreds and eventually thousands of compute nodes across space and ground. You will work on orchestration, fleet management, reliability, networking, observability, and automated deployment across highly constrained distributed environments.
Satlyt is building virtual AI data centers in space. Unlike traditional cloud infrastructure, Satlyt operates across compute nodes that may be intermittently connected, bandwidth-constrained, power-constrained, resource-constrained, and physically inaccessible.
Design and build Satlyt’s distributed compute and orchestration platform
-
Build systems for
Workload scheduling
Deployment
Monitoring
Software updates
Failure recovery
Develop fleet-management systems across heterogeneous compute environments
Build resilient infrastructure for intermittently connected nodes
Design control-plane and data-plane systems
Build networking abstractions across terrestrial and space-based links
Improve container orchestration across constrained edge hardware
Build secure remote software deployment and update mechanisms
Develop automated provisioning and configuration systems
Build observability across distributed spacecraft and ground nodes
-
Design systems that tolerate
Network partitions
High latency
Packet loss
Node failures
Stale state
Delayed communication
-
Optimize
Compute
Storage
Bandwidth
Power
Debug complex distributed systems and infrastructure failures
Build platform tooling and infrastructure automation
Turn recurring customer integration requirements into reusable platform capabilities
Work closely with the CTO on architecture decisions and technical tradeoffs
5+ years building production distributed systems, infrastructure, platform systems, or similarly complex software
Strong Linux fundamentals
-
Strong programming ability in
Go
Rust
C++
Python
-
Experience with
Kubernetes
Container runtimes
Schedulers
Distributed orchestration
Strong understanding of networking
Strong distributed systems fundamentals
Experience operating important production infrastructure
-
Experience with infrastructure operations including
Observability
CI/CD
Infrastructure automation
Configuration management
Strong ability to reason about failure modes and degraded operating conditions
Ability to make clear architecture and engineering tradeoffs
Strong ownership and execution mindset
Kubernetes internals
Edge computing
Embedded Linux
ARM-based systems
Fleet management
Distributed databases
Service meshes
Overlay networking
Satellite communications
DTN or HDTN
BPv7, CFDP, or ION
Security engineering
AI infrastructure
Robotics or autonomous systems
Remote device management
Experience operating large distributed device fleets