Python Data Engineer (Hibrido, Porto)

Join our team as a Data Engineer and help build and maintain scalable data solutions across modern data platforms.

You'll design, develop, and optimise both batch and real-time data pipelines, ensuring reliable data processing, production stability, and high-quality data delivery to support business-critical needs.

If you're passionate about data engineering, scalable architectures, and building robust production-grade solutions, we'd love to hear from you.

Key Responsibilities:

  • Design, develop, test and deploy ETL/ELT pipelines using Python and PySpark.
  • Optimise data storage and query performance across data lakes, warehouses and lakehouse formats such as Delta Lake or Iceberg.
  • Apply software-engineering best practices, including maintainable code, reviews, CI/CD and unit testing.
  • Implement monitoring, logging and alerting and resolve pipeline failures.
  • Partner with data scientists, analysts and product stakeholders and optimise Spark workloads for performance and cost.

Requisitos mínimos

  • Advanced Python, including object-oriented programming, pandas and testing frameworks.
  • Strong practical Apache Spark / PySpark experience.
  • Advanced SQL and relational databases such as PostgreSQL or MySQL.
  • Experience with a data warehouse such as Snowflake, BigQuery or Redshift.
  • Experience with at least one major cloud platform: AWS, GCP or Azure.
  • Workflow orchestration using Airflow, Prefect or Dagster.
  • Excellent written and spoken English.