Databricks SME – Operations & Administration

Contract
Remote within US
Undisclosed

Job description

Job Title: Databricks SME – Operations & Administration

Role Type: Contract

Work Arrangement: Remote

Job Summary

We are seeking an experienced Databricks Subject Matter Expert (SME) with strong hands-on experience in Databricks platform administration, production support, troubleshooting, configuration, and operational activities. The ideal candidate will be responsible for maintaining the stability, performance, availability, and security of the Databricks environment and supporting critical production workloads.

Key Responsibilities

  • Provide L2/L3 operational support and administration for enterprise Databricks environments.

  • Perform day-to-day Databricks platform administration, configuration, and troubleshooting.

  • Manage and troubleshoot Databricks Workspaces, Clusters, Jobs, Notebooks, Workflows, and SQL Warehouses.

  • Monitor production workloads and proactively identify and resolve platform and job-related issues.

  • Troubleshoot cluster failures, job failures, performance issues, connectivity issues, and resource-related problems.

  • Configure and manage cluster policies, compute resources, runtime versions, libraries, and access controls.

  • Manage Databricks Jobs/Workflows, schedules, dependencies, alerts, and failed executions.

  • Perform root cause analysis (RCA) for recurring production issues and implement permanent fixes.

  • Monitor platform utilization, cluster performance, job execution, and resource consumption.

  • Support capacity planning, performance tuning, and optimization of Databricks environments.

  • Administer users, groups, permissions, service principals, tokens, and workspace access.

  • Implement and troubleshoot RBAC and security configurations in Databricks.

  • Support integration with Azure/AWS cloud services, storage, networking, and enterprise data platforms.

  • Troubleshoot connectivity between Databricks and external systems such as ADLS/S3, databases, APIs, data warehouses, and messaging platforms.

  • Support Databricks Runtime upgrades, patches, configuration changes, and platform maintenance.

  • Coordinate with Cloud, Data Engineering, Security, Network, and Application teams for production incidents.

  • Participate in incident, problem, and change management processes.

  • Maintain operational documentation, runbooks, knowledge articles, and troubleshooting procedures.

  • Participate in on-call/after-hours production support as required.

Required Skills

  • Strong hands-on experience with Databricks administration and production operations.

  • Excellent troubleshooting skills across Databricks platform, clusters, jobs, workflows, and connectivity.

  • Strong knowledge of Databricks Workspaces, Compute/Clusters, Jobs, Workflows, Delta Lake, and Databricks SQL.

  • Experience with cluster configuration, cluster policies, autoscaling, libraries, and runtime management.

  • Experience with Databricks security, RBAC, users, groups, service principals, and access management.

  • Strong knowledge of Apache Spark and Spark troubleshooting.

  • Experience with Python and/or PySpark, SQL, and scripting.

  • Experience with at least one major cloud platform: Azure, AWS, or GCP.

  • Strong knowledge of cloud storage such as ADLS Gen2, Amazon S3, or GCS.

  • Experience with monitoring, alerting, logging, and production incident management.

  • Strong understanding of performance tuning and capacity management.

  • Excellent communication and problem-solving skills.

Preferred / Nice to Have

  • Experience with Databricks on Azure (Azure Databricks).

  • Experience with Unity Catalog and data governance.

  • Experience with Terraform/IaC for Databricks configuration.

  • Knowledge of Azure DevOps/Git and CI/CD.

  • Experience with Databricks APIs and CLI.

  • Experience supporting large-scale enterprise production environments.

  • Knowledge of ITIL, incident management, change management, and RCA processes.

Ideal Candidate: A hands-on Databricks Operations/Platform SME who can independently administer, configure, monitor, troubleshoot, optimize, and support Databricks production environments, rather than a candidate focused primarily on Databricks data engineering/development.

More information

Job skills

Databricks Administration

Cloud Infrastructure

Data Pipeline Management

SQL Performance Tuning

User Access Control

Monitoring and Reporting

Troubleshooting

Data Governance

Client company information

The client company is confidential. Details will be shared after mutual interest is confirmed.

Company overview

company-logo
Tech3pillars technologies LLC

IT Staffing & IT Consulting

For over 8 years, T3Pillars Technologies has been at the forefront of the digital transformation wave in the US market. What started as a focused IT staffing agency in the DC corridor has evolved into a comprehensive consulting firm specializing in enterprise Salesforce implementations, CRM/ERP integrations, custom application development, and managed support. Headquartered at 11974 Grey Wing Ct, Reston, VA, we are strategically positioned in the heart of the DC technology corridor — one of the most competitive technology markets in the country. This location has shaped our culture: we operate with the precision and accountability demanded by the government contractors, healthcare systems, and financial institutions that form our client base. Today, T3Pillars serves clients nationwide, bringing world-class talent and technical architecture to industries that demand reliability, compliance, and measurable outcomes. Our 50+ certified consultants have collectively delivered over 200 projects — on time, on budget, and to lasting effect.