本文へスキップ

Lead Site Reliability Engineer, Electronic Colo Trading

JPMorganChaseJapanインフラSRE
GoLinuxPython

Job Responsibilities

  • Define and champion non-functional requirements (NFRs) and availability targets for services/products in electronic colo trading.
  • Ensure reliability, performance, and operational excellence of Linux-based compute platforms; emphasize automation, standardization, and disciplined incident management.
  • Collaborate with network, trading technology, and data-center teams to optimize latency, throughput, and stability (kernel/IRQ/CPU isolation, NUMA, NIC tuning).
  • Develop and track service level indicators/SLIs and service level objectives with stakeholders; establish and manage error budgets.
  • Operate and improve observability across the stack (metrics, logs, alerting); translate signals into actionable runbooks and improvements.
  • Lead major incidents (triage, mitigation, root-cause analysis, corrective actions) and coordinate hardware-adjacent colo tasks.
  • Leverage AI capabilities to accelerate triage and post-incident analysis; validate outputs and ensure security/compliance; mentor engineers.

Technical Stack

必須スキル

  • Bachelor’s degree and 5+ years in Site Reliability Engineering or related field; formal SRE training/certification
  • Production Linux administration (RHEL/Debian/Ubuntu); strong OS, hardware, and basic networking troubleshooting
  • Linux internals, performance tuning, and diagnosing complex latency/reliability issues
  • Automation with Bash and Python or Go; configuration management, CI/CD, monitoring, observability
  • Incident management, root-cause analysis, and postmortems; drive durable improvements
  • Use of enterprise AI capabilities with validation, guardrails, and data sensitivity awareness

歓迎スキル(該当する場合)

  • Experience with electronic trading, latency-sensitive environments, or colocation data centers
  • Kernel/CPU pinning, IRQ tuning, time synchronization, and low-latency Linux optimization
  • Infrastructure-as-code patterns and standardized “golden” server builds; fleet management and patching
  • Experience operating at scale with automation-driven patching and drift control
  • Familiarity with AI-assisted reliability workflows across SDLC (CI/CD checks, runbooks, validation)

キャリア成長観点

  • 高度なレイテンシー要求を持つ取引プラットフォームの信頼性をリードすることで、ビジネス影響を直接推進できる。
  • ハードウェア/ネットワーク連携を含むデータセンター規模の運用とLinuxパフォーマンス最適化に深く精通できる。
  • AIを活用した信頼性ワークフローの設計・統治を通じて、先進的な運用 governance を築ける。
  • Platform SRE/Reliability Architectなどのキャリアパスや、金融業界のセキュリティ・ガバナンスの経験値を高める機会が豊富。

データ取得日: 2026/9/10

関連求人