Python Data Engineer (Hibrido, Porto)
Join our team as a Data Engineer and help build and maintain scalable data solutions across modern data platforms.
You'll design, develop, and optimise both batch and real-time data pipelines, ensuring reliable data processing, production stability, and high-quality data delivery to support business-critical needs.
If you're passionate about data engineering, scalable architectures, and building robust production-grade solutions, we'd love to hear from you.
Key Responsibilities:
- Design, develop, test and deploy ETL/ELT pipelines using Python and PySpark.
- Optimise data storage and query performance across data lakes, warehouses and lakehouse formats such as Delta Lake or Iceberg.
- Apply software-engineering best practices, including maintainable code, reviews, CI/CD and unit testing.
- Implement monitoring, logging and alerting and resolve pipeline failures.
- Partner with data scientists, analysts and product stakeholders and optimise Spark workloads for performance and cost.
Requisitos mínimos
- Advanced Python, including object-oriented programming, pandas and testing frameworks.
- Strong practical Apache Spark / PySpark experience.
- Advanced SQL and relational databases such as PostgreSQL or MySQL.
- Experience with a data warehouse such as Snowflake, BigQuery or Redshift.
- Experience with at least one major cloud platform: AWS, GCP or Azure.
- Workflow orchestration using Airflow, Prefect or Dagster.
- Excellent written and spoken English.