TradeStation

Site Reliability Engineer

TradeStation

Financial Services · 501-1,000 employees

5 h ago
Remote sre Mid (2-5 yrs) Full-time
Create a free account to apply — email only, no card. You can also save this posting or score it against your profile with AI.

About the role

The Site Reliability Engineer will build and maintain scalable AWS cloud infrastructure and Kubernetes clusters for the Order Execution team. They will also participate in on-call rotations, manage incident responses, and implement observability tools to ensure system reliability.

What they look for

Kubernetes AWS Terraform GitLab CI Python Golang OpenTelemetry Kafka Site Reliability Engineering Infrastructure as Code Cloud Infrastructure Monitoring Disaster Recovery Performance Tuning Linux Agile

Requirements

Candidates must have at least 3 years of experience in SRE, DevOps, or cloud infrastructure and hold a bachelor's degree in a technical field. Proficiency in Kubernetes, AWS, and infrastructure-as-code tools is required, along with the ability to work during US market hours.

Benefits

Competitive salaries Yearly bonus Comprehensive benefits Unlimited paid time off Flexible working environment TradeStation account employee benefits Trading education materials

Full description

#WeAreTradeStation

 

Who We Are:

 

TradeStation is the home of those born to trade. As an online brokerage firm and trading ecosystem, we are focused on delivering the ultimate trading experience for active traders and institutions. We continuously push the boundaries of what's possible, encourage out-of-the-box thinking, and relentlessly search for like-minded innovators.

 

At TradeStation, we are building an AI-First culture. We expect team members to embrace AI as a core part of their daily workflow, whether that’s using AI to accelerate development, enhance decision-making, improve client outcomes, or streamline internal processes. We hire, grow, and promote people who can harness AI responsibly and creatively. We treat AI as a partner in problem-solving, not just a tool; following our governance standards to ensure AI is used ethically, securely, and transparently. If you join us, you’re joining a culture where AI is how we work.

 

Are you ready to make yourself at home?

 

What We Are Looking For:

 

Do you enjoy tackling interesting problems with other smart people? Have you ever wondered how a major player in the Financial Technology industry operates? At TradeStation, you'll enjoy working in a supportive environment with the latest in microservices and streaming technology — all from the comfort of your home office.

 

We are seeking a Site Reliability Engineer to join the Order Execution – Streaming (OXS) team. OXS builds low-latency, scalable, and extensible services that provide access to order execution back-office data — the streams and APIs that deliver orders, positions, and balances to TradeStation and Monex clients across web, desktop, mobile, and third-party APIs.

 

You will be paired with senior SREs and a defined readiness program rather than dropped into the deep end, and you will progress from shadowing to independent ownership of infrastructure, observability, and on-call responsibilities. We are looking for a well-rounded engineer with genuine curiosity about how distributed systems fail, and the discipline to work carefully in an environment where real customer funds are at stake.

 

What You’ll Be Doing:

 

Reliability and Cloud Infrastructure

  • Understand, execute, and embody Site Reliability Engineering principles
  • Work and collaborate in building and maintaining AWS cloud infrastructure for the OXS development and quality engineering teams to utilize
  • Author and maintain infrastructure as code, with all AWS infrastructure created through Stacker and terraform, and checked into GitLab as configuration-as-code
  • Administer and maintain Kubernetes (EKS) clusters and the OXS workloads that run on them — node groups and upgrades, RBAC and namespace policy, resource requests/limits and autoscaling, ingress and service networking, and deployment manifests and Helm/pipeline configuration
  • Build and exercise cross-region disaster recovery for OXS services — multi-AZ and multi-region failover, backup and restore of state stores, replication of Kafka/MSK and configuration, and periodic DR testing against defined RTO and RPO targets
  • Assist with building the necessary guardrails to keep services operational and secure
  • Assist with building templates, automation, and tooling that accelerate development and reduce operational toil
  • Work in a DevOps environment, where development teams own both the development and operational responsibilities
  • Provide technical guidance to developers and less experienced engineers on cloud-native architecture, resilience patterns, and operational readiness

Observability and Service Health

  • Instrument OXS services with OpenTelemetry, build and maintain dashboards, metrics, and alerting
  • Help define and measure service level indicators and objectives for order, position, and balance streaming, including stage-level latency budgets
  • Improve alert quality — reduce noise, standardize severity-based routing, and close monitoring gaps surfaced by incidents
  • Analyze logs, traces, and Kafka consumer lag to identify degradation before customers notice it
  • Profile and tune performance across the AWS stack — EC2/EKS instance and node sizing, container CPU/memory tuning, JVM and runtime settings, MSK and ElastiCache throughput, network and load-balancer configuration — and validate improvements through load testing against latency budgets
  • Profile and tune performance across the AWS stack — EC2/EKS instance and node sizing, container CPU/memory tuning, JVM and runtime settings, MSK and ElastiCache throughput, network and load-balancer configuration — and validate improvements through load testing against latency budgets
  • Evaluate emerging technologies and recommend improvements to platform reliability, scalability, and operational efficiency

Environments, Pipelines, and Release Support

  • Support and maintain the OXS environment state across environments
  • Maintain GitLab CI pipelines, container images, and ECR lifecycle configuration; keep builds and deployments repeatable, versioned, automated, and rollback-capable
  • Support release operations alongside developers and SDETs, including change request preparation and post-deployment validation
  • Follow TradeStation change management across different environments up to production, so that changes carry a complete audit trail of approvals

Incident Response and Production Support

  • Participate in the OXS on-call rotation, progressing from shadowing to Secondary and then Primary as you complete the team's on-call readiness program
  • Review dashboards and alerts at the start of the trading day and confirm system readiness ahead of market open
  • Triage and respond to incidents using OXS runbooks, escalating early rather than late
  • Contribute to blameless postmortems, root cause analysis, and the preventive-measure work items that follow them
  • Author and improve runbooks, SOPs, and operational documentation so that knowledge does not live in one person's head

Security and Cost

  • Help remediate cloud security findings identified through Wiz in coordination with the security team
  • Support cloud cost visibility and optimization for OXS AWS accounts, flagging spend anomalies and helping right-size infrastructure

The Skills You Bring:

  • Knowledge of Kubernetes administration — cluster and node group operations, upgrades, RBAC, scheduling and resource management, autoscaling, and troubleshooting workloads and control-plane issues
  • Knowledge and proficiency in one or more modern general-purpose programming languages (e.g. Python, C#, Golang, Bash)
  • Knowledge of AWS cloud infrastructure and core services (EKS, EC2, S3, ECR, IAM, MSK, ElastiCache, RDS)
  • Experience building or modifying cloud infrastructure as code (e.g. Stacker, CloudFormation, Terraform)
  • Experience with Continuous Integration tools, GitLab CI preferred (e.g. GitLab CI, Azure DevOps, Jenkins)
  • Familiarity with observability tooling and concepts — metrics, logs, traces, dashboards, and alerting (e.g. OpenTelemetry, Grafana, Datadog, Prometheus)
  • Excellent written and verbal communication skills, with the ability to write clear runbooks and incident updates and to assist developers in building cloud native applications
  • Familiarity working in an Agile environment and demonstrated success with structured testing practices such as automated unit testing, integration testing, TDD, and continuous delivery
  • Willingness to participate in an on-call rotation supporting a production trading system, and to be responsive during US equity and futures market hours
  • Demonstrated working knowledge of a cloud provider, containers, and a CI/CD pipeline, whether from professional experience, internships, or substantial personal projects
  • Experience operating and administering Kubernetes in production, including cluster upgrades, capacity and autoscaling decisions, and troubleshooting workloads under load
  • Experience participating in disaster recovery exercises or regional failover events for a production system
  • Understanding of high-availability and disaster recovery design in AWS — multi-AZ and cross-region architectures, failover strategies, backup/restore, and RTO/RPO tradeoffs
  • Experience diagnosing and tuning performance in cloud environments — latency and throughput analysis, right-sizing, capacity planning, and load or stress testing
  • Experience with distributed and scalable cloud architectures and techniques preferred
  • Familiarity with event streaming and messaging systems, especially Apache Kafka or AWS MSK, preferred
  • Familiarity with Redis, or another distributed key/state store, preferred
  • Exposure to Microsoft SQL Server administration or migration work preferred
  • Experience managing Linux deployments in the cloud preferred
  • Experience securing cloud deployments preferred
  • Advanced understanding of network and internet technologies (e.g. DNS, TLS, TCP, UDP, HTTP, gRPC, WebSockets) preferred
  • Familiarity with change management or audited release processes (e.g. ServiceNow) preferred
  • Awareness of cloud cost management and optimization practices preferred
  • Comfort using AI tooling as part of daily engineering workflow — code assistants, and AI-assisted investigation, documentation, and reporting — preferred
  • Brokerage/trading domain knowledge and experience preferred
  • Experience with distributed and scalable cloud architectures and techniques preferred
  • Prior experience supporting a production system with defined availability or latency targets preferred
  • Experience operating Kubernetes in production, including troubleshooting workloads under load preferred
  • Exposure to Kafka topics, partitions, consumer groups, and lag troubleshooting preferred
  • Experience in financial services, brokerage, trading, or another regulated, high-availability domain preferred

Minimum Qualifications:

  • Bachelor's degree in Computer Science, Engineering, or a related technical field, or equivalent work experience
  • 3+ years of hands-on experience in site reliability engineering, DevOps, cloud infrastructure, or application development
  • Ability to work core hours aligned to US market hours (approximately 9:00 AM – 5:00 PM ET) and to participate in an on-call rotation

Desired Qualifications:

  • 2+ years of professional SRE, DevOps, or cloud infrastructure experience
  • AWS certification (e.g. Cloud Practitioner, Solutions Architect Associate, DevOps Engineer)
  • Certified Kubernetes Administrator (CKA) or equivalent hands-on Kubernetes credential
  • Professional working proficiency in English and Spanish

What We Offer:

  • Collaborative work environment
  • Competitive Salaries
  • Yearly bonus
  • Comprehensive benefits for you and your family starting Day 1
  • Unlimited Paid Time Off
  • Flexible working environment
  • TradeStation Account employee benefits, as well as full access to trading education materials

Similar roles