Speaker
Abstract
At GEICO we are on a journey to entirely modernize our Infrastructure. We are building an open-source, cloud-agnostic hybrid stack to run across public and on prem private cloud infrastructure without having to expose vendor specific stacks to our application developers. This hybrid stack gives us flexibility to run workloads wherever we need them, and to migrate significant workloads from the public cloud to our on-prem infrastructure where cost or latency are better served for those workloads.
Through that process we had to select new colocation facilities (moving from 6 facilities to 3 better balanced and geo-distributed sites), Open Hardware servers (based on the workload characteristics of our legacy and cloud footprints leveraging OpenBMC and Redfish for management), Open Network solutions (switches, routers and our own NOS for those systems), and OpenStack (including Ceph for SDS) to deliver fleet management solutions across our on prem footprint.
This change is driving 30% to 3x cost savings per workload relative to the equivalent capacity, latency, and up time in our current cloud providers. We have also completely redesigned our current on prem network and servers from a demilitarized zone isolated network approach on MPLS cirtuits to a fully untrusted network (only decrypt where the user/account/application is allowed to have access) using direct internet access, and profoundly simplifying our hardware skus (going from over 200 instances in the public cloud down to 5 primary, and 15 specialty solutions to be phased out as our applications modernize).
In this session we will walk through the hardware selection process taking our workload characteristics from the cloud and using that to optimize a subset of SKUs for our on prem cloud.
Interview
I run infrastructure engineering for GEICO, which includes our hardware systems (compute, storage, networking, AI, etc.), our workflow automation, provisioning, and fleet management tools for the physical assets, and our full hybrid cloud stack (data protection services, identity and access management tools, OS, runtime, and container management solutions, cluster management, and service mesh across our public and private cloud footprint).
Making it easier for developers to decode public cloud instances into a physical footprint, helping demystify where private cloud can be more efficient and where public cloud is optimal.
Devops folks trying to understand the tradeoffs between public and private cloud for overall reliability, security, and efficiency.
More understanding of why private cloud is becoming increasingly necessary WITH public cloud offerings for Enterprise institutions. Where the cloud isn’t serving customers well. How to create a footprint that meets the needs of an actual business.
I hate the word disruption: it feels like a buzz word. I believe the pendulum has swung to where AI and data security are requiring a hybrid approach to Infrastructure, and I’m looking to the open source community to come together to create the right design patterns to ensure we are able to run hybrid cloud efficiently and effectively.
Topics
QCon San Francisco 2024 is a three day conference for senior software engineers, architects and team leads. An international program committee of working engineers selects every session. Patterns and practices, not products and pitches.
From the same track
Wednesday 20 November
10:35 Seacliff ABC Session AI/ML Unleashing Llama's Potential: CPU-Based Fine-Tuning Anil Rajput, Rema Hariharan Generative AI landscape is rapidly changing as new models are appearing in horizon every few days. However, the hardware and software characteristics of these models have many similar patterns and execution phases. 11:45 Seacliff ABC Session AI HW/SW optimization Maximizing Deep Learning Performance on CPUs using Modern Architectures Bibek Bhattarai AI Technical Lead @Intel, Computer Scientist Invested in Hardware-Software Optimization, Building Scalable Data Analytics, Mining, and Learning Systems As deep learning continues to drive advancements across various industries, efficiently navigating the landscape of specialized AI hardware has a huge impact on cost and speed of operation. 13:35 Seacliff ABC Session Hybrid cloud Evaluating and Deploying State-of-the-Art Hardware to Meet the Challenges of Modern Workloads Rebecca Weekly VP of Infrastructure @GEICO At GEICO we are on a journey to entirely modernize our Infrastructure. We are building an open-source, cloud-agnostic hybrid stack to run across public and on prem private cloud infrastructure without having to expose vendor specific stacks to our application developers. 14:45 Seacliff ABC Session RISC-V Optimizing Custom Workloads with RISC-V Ludovic Henry Member of Technical Staff @Rivos, Performance-Minded Engineer, Hardware & Software, Previously @Xamarin, @Microsoft, @Datadog This talk will explore how RISC-V architecture can accelerate custom workloads, focusing on AI/ML applications. We’ll start by examining the RISC-V ecosystem and its increasing relevance in the software development landscape. 15:55 Seacliff ABC Session High-Resolution Platform Observability Brian Martin Co-founder and Software Engineer @IOP Systems, Focused on High-Performance Software and Systems, Previously @Twitter Many observability tools fail to provide us with the relevant insights for understanding hardware health and utilization.