Staff AI-Ops Engineer
Charlotte, NC, US, 28202
SMBC Group is a top-tier global financial group. Headquartered in Tokyo and with a 400-year history, SMBC Group offers a diverse range of financial services, including banking, leasing, securities, credit cards, and consumer finance. The Group has more than 130 offices and 80,000 employees worldwide in nearly 40 countries. Sumitomo Mitsui Financial Group, Inc. (SMFG) is the holding company of SMBC Group, which is one of the three largest banking groups in Japan. SMFG’s shares trade on the Tokyo, Nagoya, and New York (NYSE: SMFG) stock exchanges.
In the Americas, SMBC Group has a presence in the US, Canada, Mexico, Brazil, Chile, Colombia, and Peru. Backed by the capital strength of SMBC Group and the value of its relationships in Asia, the Group offers a range of commercial and investment banking services to its corporate, institutional, and municipal clients. It connects a diverse client base to local markets and the organization’s extensive global network. The Group’s operating companies in the Americas include Sumitomo Mitsui Banking Corp. (SMBC), SMBC Nikko Securities America, Inc., SMBC Capital Markets, Inc., SMBC MANUBANK, JRI America, Inc., SMBC Leasing and Finance, Inc., Banco Sumitomo Mitsui Brasileiro S.A., and Sumitomo Mitsui Finance and Leasing Co., Ltd.
Role Description
As a Staff AI-Ops Engineer in the Platform Engineering team, you will play a pivotal role in operationalizing, monitoring, and governing the AI/GenAI platform and the models, pipelines, and agents that run on it. You will work closely with stakeholders in architecture, technology, data and business organizations. You will partner with Azure, Databricks, and other infrastructure providers to build and operate the MLOps/LLMOps backbone of the AI/GenAI platform, ensuring reliable, observable, secure, and cost-efficient AI systems in production. To succeed in this role, you should be a fast learner who can quickly adopt upcoming AI/GenAI operational tooling, and an accomplished coder capable of building enterprise-scale automation, CI/CD, and observability systems.
This is a unique opportunity to own the operational excellence of the GenAI technology stack—bridging the gap between one-off experiments and production-grade AI systems—ensuring industrial-grade reliability, compliance, and efficiency in a high-stakes financial environment.
Role Objectives
- Operationalize the AI Platform: Design and operate the MLOps/LLMOps backbone for the AI platform on Databricks and Azure Cloud Services, standardizing how models, prompts, pipelines, and agents are built, promoted, and run.
- Build CI/CD and release engineering: Develop automated CI/CD pipelines and infrastructure-as-code for models, prompts, and agents using Databricks Asset Bundles across DEV/QA/REL/PROD, with canary, blue/green, shadow, and automated-rollback deployment strategies.
- Own governance, versioning and auditability: Implement end-to-end lineage and version control across data, prompts, retrievals, models, and responses using MLflow (Prompt Registry, Tracing, Experiments/Runs), delivering audit-ready artifacts and enforceable quality gates for internal and regulatory review.
- Monitoring, drift and cost governance: Build observability for data quality, data and model drift, retrieval and hallucination/grounding health, application performance, and business KPIs, with cost visibility, inference optimization, and FinOps-aligned governance.
- Testing, evaluation and validation: Establish automated regression, A/B, canary, shadow, and champion-challenger validation with golden datasets, evaluation rubrics, and human-in-the-loop review to certify quality and safety before and after release.
- Responsible AI and security controls: Operationalize responsible-AI guardrails (bias/harm detection, explainability, safety) and security/privacy controls—authentication, authorization, secret management, and protection of data, models, prompts, and embeddings—across the inference and agent-tool layers.
- Operational readiness and run management: Support reliable day-2 operations through model cards, API/SLA contracts, runbooks, incident response, and escalation readiness.
- Evaluate emerging technology: Proactively identify and evaluate emerging AI-Ops tooling and integrate those that improve reliability, observability, and cost efficiency.
- Technical mentorship: Mentor and educate broader engineering teams on MLOps/LLMOps best practices and platform operational capabilities.
Qualifications and Skills
- Bachelor’s degree in Computer Science, Machine Learning, Data Science, or related field.
- 5+ years of hands-on experience deploying, operating, and maintaining GenAI or advanced ML models in production environments, with a strong focus on MLOps/LLMOps.
- 3+ years of experience in Python and GenAI frameworks/tools e.g. Databricks Vector Search, Azure AI Search, Azure AI document intelligence, LangGraph, haystack, Llama Index etc.
- Deep, hands-on expertise with MLOps/LLMOps tooling (e.g. MLflow — Prompt Registry, Tracing, Experiments, Model Serving), data platforms (e.g. Databricks, Databricks Asset Bundles) and cloud platforms (e.g. Azure).
- Demonstrated experience developing and deploying RESTful services, containerization, and automated CI/CD systems.
- Proven experience building observability, monitoring, and alerting for AI systems — data and model drift, evaluation metrics, hallucination/grounding health, performance, and cost (FinOps).
- Working knowledge of prompt engineering, embedding models, RAG evaluation, and vector databases sufficient to instrument, test, and monitor GenAI applications.
- Working knowledge of ML libraries e.g. PyTorch, TensorFlow, Hugging Face Transformers.
- Familiarity with AI governance, responsible-AI, and security/privacy controls in a regulated (e.g. financial services) environment.
- Excellent communication and collaboration skills; proven ability to influence and partner with technical and non-technical stakeholders.
SMBC’s employees participate in a Hybrid workforce model that provides employees with an opportunity to work from home, as well as, from an SMBC office. SMBC requires that employees live within a reasonable commuting distance of their office location. Prospective candidates will learn more about their specific hybrid work schedule during their interview process. Hybrid work may not be permitted for certain roles, including, for example, certain FINRA-registered roles for which in-office attendance for the entire workweek is required.
SMBC provides reasonable accommodations during candidacy for applicants with disabilities consistent with applicable federal, state, and local law. If you need a reasonable accommodation during the application process, please let us know at accommodations@smbcgroup.com.
Nearest Major Market: Charlotte