Job Description
Data Engineer is a individual contributor (IC) responsible for designing, building, and evolving the companys data platform and high-impact data products so analytics, BI, AI/ML, and operational use cases are reliable, secure, and scalable. This role blends hands-on engineering with technical leadership: setting patterns and standards, driving cross-team alignment, and unblocking complex delivery across the data ecosystem.
Roles & Responsibilities:
- Shape and drive enterprise-wide data architecture strategy: Define and evolve the long-term technical vision for scalable, resilient data infrastructure across multiple business units and domains.
- Lead large-scale, cross-functional initiatives: Architect and guide the implementation of data platforms and pipelines that enable analytics, AI/ML, and BI at an organizational scale.
- Pioneer advanced and forward-looking solutions: Introduce novel approaches in real-time processing, hybrid/multi-cloud, and AI/ML integration to transform how data is processed and leveraged across the enterprise.
- Mentor and develop senior technical leaders: Influence Principal Engineers, Engineering Managers, and other Staff Engineers; create a culture of deep technical excellence and innovation.
- Establish cross-org technical standards: Define and enforce best practices for data modeling, pipeline architecture, governance, and compliance at scale.
- Solve the most complex, ambiguous challenges: Tackle systemic issues in data scalability, interoperability, and performance that impact multiple teams or the enterprise as a whole.
- Serve as a strategic advisor to executive leadership: Provide technical insights to senior executives on data strategy, emerging technologies, and long-term investments.
- Represent the organization as a thought leader: Speak at industry events/conferences, publish thought leadership, contribute to open source and standards bodies, and lead partnerships with external research or academic institutions.
Technical Skills:
- 4+ years of experience
- Mastery of GCP Data Ecosystem: Deep authority in architecting, designing, and building complex solutions using BigQuery, Dataflow, Dataproc, and Pub/Sub for massive-scale batch and streaming workloads
- Advanced programming and infrastructure capabilities: Expertise in Python, or Java, along with infrastructure-as-code tools like Terraform or Cloud Deployment Manager.
- Leadership in streaming and big data systems: Authority in tools such as BigQuery, Dataflow, Dataproc, Pub/sub for both batch and streaming workloads.
- Enterprise-grade governance and compliance expertise: Design and implement standards for data quality, lineage, security, privacy (e.g., GDPR, HIPAA), and auditability across the organization.
- Build and maintain curated datasets and semantic layers that meet defined contracts (freshness, accuracy, completeness, schema stability).
- Implement scalable data storage patterns (warehouse/lakehouse, partitioning, clustering, file layout, table formats) and performance optimization.
- Engineer secure data access patterns (RBAC/ABAC, row/column-level security, tokenization/masking, encryption, key management).
- Data Quality & Lineage: Implementation of automated data validation and lineage tracking standards to ensure "Single Source of Truth" reliability across diverse business units.
- Cost Optimization (FinOps): Advanced skill in managing cloud spend through granular resource monitoring and the implementation of cost-efficient data lifecycle policies.
- Strategic integration with AI/ML ecosystems: Architect platforms that serve advanced analytics and AI workloads (Vertex AI, TFX, MLflow).
- Exceptional ability to influence across all levels: Communicate technical vision to engineers, influence strategic direction with executives, and drive alignment across diverse stakeholders.
- Recognized industry leader: Demonstrated track record through conference presentations, publications, open-source contributions, or standards development.
Must Have Skills:
- Deep expertise in data architecture, distributed systems, and multi-cloud (GCP, AWS, Azure)
- Python or Java, infrastructure-as-code (e.g. Terraform)
- Big data tools: BigQuery, Dataflow, Dataproc, Pub/Sub (batch + streaming)
- Data governance, privacy, and compliance (e.g. GDPR, HIPAA)
- Systemic Reliability: Shifting from fixing individual bugs to preventing entire classes of incidents through systematic reliability engineering and observability.
About this job listing
This job opportunity is provided through our
external job listing network. MyJobAlerts helps
you discover job opportunities and redirects you
to the original listing to apply.