Principal AI Data Engineer

Austin, TX, US Senior Data Engineer

Interested in this Data Engineer role at Presidio?

Apply Now →

Skills & Technologies

AzurePower BiPythonRagSalesforceTableau

About This Role

AI job market dashboard showing open roles by category

Presidio, Where Teamwork and Innovation Shape the Future

At Presidio, we're at the forefront of a global technology revolution, transforming industries through cutting\-edge digital solutions and next\-generation AI. We empower businesses \- and their internal customers \- to achieve more through innovation, automation, and intelligent insights.

The Role

Responsibilities Include:

Technical Leadership

  • Establish engineering standards, development practices, and implementation patterns for enterprise AI and data platform solutions.
  • Mentor engineers through architecture reviews, code reviews, technical coaching, and engineering best practices.
  • Evaluate emerging technologies and recommend improvements to the enterprise AI and data platform.
  • Partner with the AI Data Architect to translate enterprise strategy into scalable, secure, and production\-ready technical solutions.
  • Promote engineering excellence across reliability, maintainability, automation, and operational support.
  • Provide technical leadership in evaluating implementation trade\-offs and recommend improvements that strengthen the enterprise architecture while maintaining alignment with strategic objectives.

Data Platform Engineering (Microsoft Fabric \& Azure)

  • Build and operate the enterprise lakehouse on Microsoft Fabric and Microsoft Azure, implementing the domain\-oriented data products, medallion\-layer structures, and Fabric\-based semantic models defined in the enterprise architecture.
  • Develop, test, and maintain data pipelines for ingestion, transformation, and serving using Fabric\-native tooling, Python, Spark, and SQL, with automated data validation to ensure integrity and timeliness.
  • Administer the Fabric and Azure data environments: capacity, workspaces, deployment pipelines, monitoring, and cost management.
  • Own performance tuning and operational excellence for the data platform, including incident response, root\-cause analysis, and continuous improvement.
  • Establish and maintain engineering practices for the platform: version control, CI/CD, code review, testing standards, and release management.

Semantic Model \& Data Product Implementation

  • Implement enterprise semantic models and certified data products to specification, encoding governed metric definitions, calculation logic, and business context from the metrics registry.
  • Implement row\-level and object\-level security in Fabric and OneLake that mirrors source\-system permissions (e.g., Salesforce roles and visibility rules) to protect sensitive pipeline, customer, and people data.
  • Integrate source systems — CRM (Salesforce), CPQ, PSA, ERP, HRIS, and finance platforms — into the enterprise model so revenue, pipeline, people, cost, and customer data are consistently defined and analytics\-ready.
  • Modernize data flows from legacy and server\-based applications into the lakehouse, with reconciliation and validation frameworks that prove parity between legacy outputs and modernized models.
  • Connect governed, certified data sources to Data Visualization Platforms (e.g., Power BI, Tableau) and partner with BI developers to migrate duplicated logic into shared enterprise models.

AI Solution Engineering

  • Build the retrieval and grounding infrastructure — semantic model endpoints, metadata services, RAG patterns, certified MCP connectors, and context APIs — that lets AI applications and agents answer business questions with governed data.
  • Engineer the enterprise context layer in partnership with the AI Data Architect and AI Enablement function, making curated business context, policies, and definitions available to AI tools.
  • Implement guardrails, access controls, and quality gates for AI data consumption in accordance with company policies.

Data Quality \& Operations

  • Implement automated data quality frameworks: validation rules, anomaly detection, reconciliation checks, and monitoring aligned to established quality standards.
  • Maintain lineage, documentation, and metadata for pipelines, models, and data products to support governance, certification, and auditability.
  • Support current\-state assessment and knowledge capture from existing systems, prior development efforts, and third\-party contractors, converting institutional knowledge into documented, maintainable code.

Collaboration

  • Partner daily with the AI Data Architect to refine designs based on implementation realities, propose technical alternatives, and deliver iteratively.
  • Work with BI developers, analysts, and domain teams to gather technical requirements and deliver reliable, well\-documented data products.
  • Mentor and review the work of internal engineers and contractors, raising the engineering bar across the data function.

Required Skills and Experience:

  • Bachelor's degree in Computer Science, Information Systems, Data Engineering, or a related field, or equivalent practical experience.
  • 10\+ years progressive experience in data engineering, software engineering, cloud data platforms, or enterprise analytics engineering. 5\+ years of hands\-on experience designing, building, and operating enterprise data platforms using Microsoft Azure, Microsoft Fabric, Databricks, Snowflake, or comparable data technologies.
  • Significant hands\-on experience building and operating enterprise data platforms in production, including lakehouse and medallion architectures, domain\-oriented data products, and semantic models.
  • Deep, hands\-on expertise with Microsoft Fabric and Microsoft Azure data services: lakehouse, Data Factory, notebooks, semantic models, deployment pipelines, and identity\-based access control (e.g., Entra ID).
  • Proven experience consolidating heterogeneous legacy source systems (e.g., mainframe, Oracle, PostgreSQL, on\-premises SQL Server) into modern cloud data platforms, including reconciliation and validation across migrations.
  • Strong programming skills in Python/PySpark and SQL, with experience engineering high\-volume production pipelines with automated auditing, validation, and recovery patterns.
  • Hands\-on Salesforce experience, including administration and integration of SFDC data and permission models into analytical platforms.
  • Experience implementing row\-level security and access controls in analytics platforms that mirror source\-system permission models.
  • Demonstrated engineering discipline: version control, structured deployment (e.g., Fabric deployment pipelines), testing, and production support.
  • Strong communication skills and the ability to work effectively with architects, analysts, business stakeholders, and third\-party contractors.

Preferred Skills and Professional Experience:

  • Experience building AI\-ready data foundations: RAG pipelines, vector/semantic retrieval, MCP or similar connector frameworks, or agent\-based data access patterns.
  • Experience with Power BI semantic model development (DAX, M, Tabular Editor) and/or Tableau connectivity and certified data sources.
  • Experience in sales operations, revenue operations, or go\-to\-market analytics domains, including territory, pipeline, and quota data models.
  • Experience with data quality tooling, observability, and automated reconciliation frameworks.
  • Experience working in contractor\-heavy or transition environments, including knowledge capture, code remediation, and acquisition data integration.
  • Relevant certifications (e.g., Microsoft Fabric, Azure Data Engineer, Salesforce).

Technical Skills Snapshot

  • Cloud \& Platform: Microsoft Fabric (OneLake, lakehouse, Direct Lake), Microsoft Azure data \& analytics services, Data Factory.
  • Engineering: Python/PySpark, SQL, notebook\-based ETL/ELT, version control, Fabric deployment pipelines, automated validation.
  • Modeling: Semantic and dimensional modeling (star/snowflake), medallion architecture, data product implementation, DAX/M.
  • Legacy Modernization: Mainframe, Oracle, PostgreSQL, and SQL Server consolidation into cloud lakehouse platforms.
  • Business Systems: Salesforce/CRM administration and integration, CPQ, PSA, ERP, HRIS.
  • Analytics \& BI: Power BI and/or Tableau connectivity, certified data sources, row\-level security.
  • AI Engineering: RAG, metadata/context services, MCP connectors, AI data guardrails.

Your future at Presidio

JoiningPresidio means stepping into a culture of trailblazers \- thinkers, builders, and collaborators \- who push the boundaries of what's possible. With our expertise AI\-driven analytics, cloud solutions, cybersecurity, and next\-gen infrastructure, we enable businesses to stay ahead in an ever\-evolving digital world.

Here, your impact is real. Whether you're harnessing the power of Generative AI, architecting resilient digital ecosystems, or driving data\-driven transformation, you'll be part of a team that is shaping the future.

Ready to innovate? Let's redefine what's next\-together.

About Presidio

Presidio is committed to hiring the most qualified candidates to join our amazing culture. We aim to attract and hire top talent from all backgrounds, including underrepresented and marginalized communities. We encourage women, people of color, people with disabilities, and veterans to apply for open roles at Presidio. Diversity of skills and thought is a key component to our business success.

At Presidio, speed and quality meet technology and innovation. Presidio is a trusted ally for organizations across industries with a decades\-long history of building traditional IT foundations and deep expertise in AI and automation, security, networking, digital transformation, and cloud computing. Presidio fills gaps, removes hurdles, optimizes costs, and reduces risk. Presidio's expert technical team develops custom applications, provides managed services, and enables actionable data insights and builds forward\-thinking solutions that drive strategic outcomes for clients globally. For more information visit

*Applications will be accepted on a rolling basis.*

*Presidio has a strong commitment to the community we serve and our employees. As an Equal Opportunity Employer, we strive to have a workforce that includes the community we serve.*

*Presidio is an Equal Opportunity Employer Disability/Vets. We evaluate qualified applicants without regard to race, color, religion, sex, age, national origin, disability, veteran status, genetic information, and other legally protected categories.*

*The "Know Your Rights" Poster is available here: https://www.eeoc.gov/poster*

*Presidio EEO Policy Statement is available here: https://www.presidio.com/careers*

*Presidio is committed to working with and providing reasonable accommodations to individuals with disabilities. If you need a reasonable* *accommodation* *because of a disability for any part of the employment process, please send an e\-mail to recruitment@presidio.com and let us know the nature of your request and your contact information.*

*Presidio is a VEVRAA Federal Contractor requesting priority referrals of protected veterans for its openings. State Employment Services, please provide priority referrals to.*

*Notice of Massachusetts Candidates: It is unlawful in Massachusetts to require or administer a lie detector test as a condition of employment or continued employment. An employer who violates this law shall be subject to criminal penalties and civil liability.*

*Recruitment Agencies, Please Note:* Presidio does not accept unsolicited agency resumes/CVs. Do not forward resumes/CVs to our career's email address, Presidio employees or any other means. Presidio is not responsible for any feeds related to unsolicited resumes/CVs.

\#LI\-FI

Equal Opportunity Employer/Protected Veterans/Individuals with Disabilities

This employer is required to notify all applicants of their rights pursuant to federal employment laws. For further information, please review the Know Your Rights (https://www.eeoc.gov/poster) notice from the Department of Labor.

Role Details

Company Presidio
Title Principal AI Data Engineer
Location Austin, TX, US
Category Data Engineer
Experience Senior
Salary Not disclosed
Remote No

About This Role

Data Engineers build the pipelines that feed AI models. They design ETL workflows, manage data lakes, and ensure training and inference data is clean, timely, and accessible. Without good data engineering, AI projects fail. It's that simple.

The AI era has expanded the data engineer's scope far beyond batch ETL jobs. You're building real-time embedding pipelines for RAG systems, managing vector databases, ensuring training data quality at scale, and building the infrastructure that lets ML teams iterate on data as fast as they iterate on models. Data quality is the biggest predictor of model quality, and you're the person responsible for it.

Across the 3,708 AI roles we're tracking, Data Engineer positions make up 1% of the market. At Presidio, this role fits into their broader AI and engineering organization.

Data Engineer demand in AI contexts is strong and growing. Every company building AI needs clean, reliable data pipelines. The shift toward real-time AI applications (chatbots, recommendation engines, agent systems) means data engineering is more critical than ever. Companies are willing to pay premium salaries for data engineers with AI/ML pipeline experience.

What the Work Looks Like

A typical week includes: debugging a data pipeline that's producing stale embeddings for the RAG system, optimizing a Spark job that processes training data, building a data quality monitoring dashboard, meeting with the ML team to understand their next data requirements, and writing dbt models that transform raw event data into ML-ready features. The work is deeply technical and high-impact.

Data Engineer demand in AI contexts is strong and growing. Every company building AI needs clean, reliable data pipelines. The shift toward real-time AI applications (chatbots, recommendation engines, agent systems) means data engineering is more critical than ever. Companies are willing to pay premium salaries for data engineers with AI/ML pipeline experience.

Skills Required

Azure (24% of roles) Power Bi (5% of roles) Python (51% of roles) Rag (23% of roles) Salesforce (4% of roles) Tableau (4% of roles)

SQL, Python, and distributed systems (Spark, Airflow, dbt) are core. Cloud data platforms (Snowflake, BigQuery, Redshift) are increasingly standard. Many AI-focused roles also want familiarity with vector databases and embedding pipelines. Understanding data modeling, pipeline orchestration, and data quality frameworks covers the essentials.

AI-specific data engineering skills include: building feature stores, managing training data versioning, implementing data lineage tracking, and building real-time embedding pipelines. Experience with streaming systems (Kafka, Flink) is valuable for real-time AI applications. Understanding ML data requirements (balanced datasets, data augmentation, evaluation set construction) makes you much more effective working with ML teams.

Strong postings specify the data stack, mention ML pipeline work, and describe the scale of data you'll be working with. Look for companies that understand the connection between data quality and model quality. Avoid roles that conflate data engineering with data analysis.

Compensation Benchmarks

Data Engineer roles pay a median of $178,800 based on 40 positions with disclosed compensation. Senior-level AI roles across all categories have a median of $230,000.

Across all AI roles, the market median is $217,500. Top-quartile compensation starts at $272,100. The 90th percentile reaches $325,000. For comparison, the highest-paying categories include AI Safety ($300,000) and Research Engineer ($280,000). By seniority level: Entry: $120,000; Mid: $200,000; Senior: $230,000; Director: $272,150; VP: $250,000.

Presidio AI Hiring

Presidio has 2 open AI roles right now. They're hiring across Data Engineer, AI/ML Engineer. Positions span Austin, TX, US, US.

Location Context

AI roles in Austin pay a median of $214,343 across 87 tracked positions.

Career Path

Common paths into Data Engineer roles include Backend Engineer, Database Administrator, Analytics Engineer.

From here, career progression typically leads toward Senior Data Engineer, ML Engineer, Data Platform Lead.

Master SQL and Python first. Then learn a distributed processing framework (Spark or its modern alternatives) and a pipeline orchestrator (Airflow, Dagster, Prefect). Build a portfolio project that demonstrates end-to-end pipeline construction: ingest, transform, validate, serve. If you want to specialize in AI data engineering, add vector databases and embedding pipelines to your skill set.

What to Expect in Interviews

Expect SQL deep-dives (query optimization, partitioning strategies, data modeling), Python coding focused on data pipeline patterns, and system design questions about building scalable ETL workflows. Companies with ML teams will ask about feature stores, embedding pipelines, and training data management. Be ready to discuss data quality monitoring, pipeline orchestration, and how you'd handle schema evolution in a production data lake.

When evaluating opportunities: Strong postings specify the data stack, mention ML pipeline work, and describe the scale of data you'll be working with. Look for companies that understand the connection between data quality and model quality. Avoid roles that conflate data engineering with data analysis.

AI Hiring Overview

The AI job market has 3,708 open positions tracked in our dataset. By seniority: 102 entry-level, 1,705 mid-level, 1,469 senior, and 432 leadership roles (Director, VP, C-Level). Remote roles make up 14% of the market (508 positions). The remaining 3,180 roles require on-site or hybrid attendance.

The market median for AI roles is $217,500. Top-quartile compensation starts at $272,100. The 90th percentile reaches $325,000. Highest-paying categories: AI Safety ($300,000 median, 21 roles); Research Engineer ($280,000 median, 147 roles); AI Architect ($254,798 median, 67 roles).

Data Engineer demand in AI contexts is strong and growing. Every company building AI needs clean, reliable data pipelines. The shift toward real-time AI applications (chatbots, recommendation engines, agent systems) means data engineering is more critical than ever. Companies are willing to pay premium salaries for data engineers with AI/ML pipeline experience.

The AI Job Market Today

The AI job market spans 3,708 open positions across 16 role categories. The largest categories by volume: AI/ML Engineer (2,605), Data Scientist (310), AI Software Engineer (259). These three account for the majority of open positions, though smaller categories often have higher per-role compensation because of specialized skill requirements.

The seniority mix tells a story about where AI teams are in their maturity. Entry-level roles (102) are outnumbered by mid-level (1,705) and senior (1,469) positions, reflecting that most companies are past the 'build a team from scratch' phase and need experienced engineers who can ship production systems. Leadership roles (Director, VP, C-Level) total 432 positions, representing the bottleneck between technical execution and organizational strategy.

Remote work availability sits at 14% of all AI roles (508 positions), with 3,180 requiring on-site or hybrid attendance. The remote share has stabilized after the post-pandemic correction. Senior and specialized roles (Research Scientist, ML Architect) are more likely to be remote-eligible than entry-level positions, partly because experienced hires have more negotiating power and partly because these roles require less hands-on mentorship.

AI compensation is structured in clear tiers. The market median sits at $217,500. Top-quartile roles start at $272,100, and the 90th percentile reaches $325,000. These figures include base salary with disclosed compensation. Total compensation (including equity, bonuses, and sign-on) runs 20-40% higher at companies that offer those components.

Category matters for compensation. AI Safety roles lead at $300,000 median, while Prompt Engineer roles sit at $140,000. The spread between highest and lowest-paying categories reflects the premium on specialized technical skills versus broader analytical roles.

The most in-demand skills across all AI postings: Python (1,890 postings), Aws (1,103 postings), Azure (877 postings), Rag (855 postings), Gcp (631 postings), Prompt Engineering (560 postings), Pytorch (545 postings), Claude (498 postings). Python dominates, appearing in the vast majority of role descriptions regardless of category. Cloud platform experience (AWS, GCP, Azure) is the second most common requirement. The newer entrants to the top skills list (RAG, vector databases, LLM APIs) reflect the shift from traditional ML toward generative AI applications.

Frequently Asked Questions

Based on 40 roles with disclosed compensation, the median salary for Data Engineer positions is $178,800. Actual compensation varies by seniority, location, and company stage.
SQL, Python, and distributed systems (Spark, Airflow, dbt) are core. Cloud data platforms (Snowflake, BigQuery, Redshift) are increasingly standard. Many AI-focused roles also want familiarity with vector databases and embedding pipelines. Understanding data modeling, pipeline orchestration, and data quality frameworks covers the essentials.
About 14% of the 3,708 AI roles we track offer remote work. Remote availability varies by company and seniority level, with senior and leadership roles more likely to offer location flexibility.
Presidio is among the companies actively hiring for AI and ML talent. Check our company profiles for detailed breakdowns of open roles, salary ranges, and hiring trends.
Common next steps from Data Engineer positions include Senior Data Engineer, ML Engineer, Data Platform Lead. Progression depends on whether you lean toward technical depth, people management, or product strategy.

Get Weekly AI Career Intelligence

Salary data, skills demand, and market signals from 16,000+ AI job postings. Every Monday.