Data Engineer
New York, NY, US, 10172
SMBC Group is a top-tier global financial group. Headquartered in Tokyo and with a 400-year history, SMBC Group offers a diverse range of financial services, including banking, leasing, securities, credit cards, and consumer finance. The Group has more than 130 offices and 80,000 employees worldwide in nearly 40 countries. Sumitomo Mitsui Financial Group, Inc. (SMFG) is the holding company of SMBC Group, which is one of the three largest banking groups in Japan. SMFG’s shares trade on the Tokyo, Nagoya, and New York (NYSE: SMFG) stock exchanges.
In the Americas, SMBC Group has a presence in the US, Canada, Mexico, Brazil, Chile, Colombia, and Peru. Backed by the capital strength of SMBC Group and the value of its relationships in Asia, the Group offers a range of commercial and investment banking services to its corporate, institutional, and municipal clients. It connects a diverse client base to local markets and the organization’s extensive global network. The Group’s operating companies in the Americas include Sumitomo Mitsui Banking Corp. (SMBC), SMBC Nikko Securities America, Inc., SMBC Capital Markets, Inc., SMBC MANUBANK, JRI America, Inc., SMBC Leasing and Finance, Inc., Banco Sumitomo Mitsui Brasileiro S.A., and Sumitomo Mitsui Finance and Leasing Co., Ltd.
The anticipated salary range for this role is between $73,000.00 and $98,000.00. The specific salary offered to an applicant will be based on their individual qualifications, experiences, and an analysis of the current compensation paid in their geography and the market for similar roles at the time of hire. The role may also be eligible for an annual discretionary incentive award. In addition to cash compensation, SMBC offers a competitive portfolio of benefits to its employees.
Role Description
We are seeking a Databricks Engineer with AWS expertise to build, optimize, and maintain our enterprise data Lakehouse infrastructure. In this role, you will be responsible for designing high-performance data pipelines, implementing regulatory and real time reporting along with advanced analytics environments, and ensuring seamless integration between Databricks and core AWS services to support real-time financial trading data consumption.
Role Objectives
- Lakehouse Architecture: Design and implement robust data pipelines using the Databricks Medallion Architecture (Bronze, Silver, Gold layers) to process structured and unstructured data.
- Pipeline Automation: Develop, scale, and orchestrate complex data workflows utilizing Databricks Jobs and Delta Live Tables (DLT).
- Data Ops & CI/CD Deployment: Standardize and automate the deployment of Databricks assets, workspace configurations, and code pipelines across Dev, QA, and Production environments.
- AWS Integration: Ensure seamless data cataloging, storage, and movement across the AWS ecosystem, specifically integrating Databricks with Amazon S3, AWS Glue, and AWS IAM for secure access control.
- Performance Optimization: Tune Spark clusters, optimize Delta Lake storage (e.g., Z-Ordering, partitioning), and manage compute costs within the AWS environment.
- Data Governance & Security: Implement fine-grained data access controls, data lineage, and auditing using Unity Catalog or native cloud security controls.
- Collaboration: Partner with Data Scientists, Risk Managers, and downstream analytics teams to deliver clean, business-ready data views for reporting and AI modeling.
Qualifications and Skills
- Professional experience in architectural design and development within the Databricks platform, working in an AWS cloud environment.
- CI/CD & DevOps Tooling: Proven proficiency in automated deployments using Databricks Asset Bundles (DABs), Terraform (specifically the Databricks and AWS providers), and standard Git pipelines (e.g., GitHub Actions, GitLab CI/CD, or AWS CodePipeline).
- Technical Proficiency: Programming skills in Python (PySpark) and SQL for complex data manipulation and transformation.
- Core Concepts: Strong understanding of Apache Spark internals, Delta Lake mechanics, and streaming data concepts (e.g., interacting with Amazon MSK or Kafka data streams).
- Data Engineering Stack: Proven experience building production-grade ETL/ELT pipelines, handling data schema validation, and cleansing raw capture feeds.
- Certifications: Databricks Certified Data Engineer Professional or AWS Certified Data Engineer – Professional is highly advantageous.
Additional Requirements
SMBC’s employees participate in a Hybrid workforce model that provides employees with an opportunity to work from home, as well as, from an SMBC office. SMBC requires that employees live within a reasonable commuting distance of their office location. Prospective candidates will learn more about their specific hybrid work schedule during their interview process. Hybrid work may not be permitted for certain roles, including, for example, certain FINRA-registered roles for which in-office attendance for the entire workweek is required.
SMBC provides reasonable accommodations during candidacy for applicants with disabilities consistent with applicable federal, state, and local law. If you need a reasonable accommodation during the application process, please let us know at accommodations@smbcgroup.com.
Nearest Major Market: New York City