Implement ETL/ELT solutions and data integration between multiple systems and data sources.
Design, implement, and maintain data pipelines to ingest, store, and process large volumes of data.
Collaborate with other teams to ensure data security and compliance.
Design, develop, and implement solutions to optimize the performance and scalability of data processing systems.
Requirements and qualifications:
Bachelor's degree in Computer Science, Computer Engineering, Information Systems, Systems Analysis and Development, or similar;
Intermediate English;
Apply knowledge of data pipeline concepts and tools to implement data transformation, cleaning, and aggregation tasks.
Use data pipeline frameworks and libraries to automate data processing tasks and optimize workflows.
Follow best practices for developing data pipelines, including version control, testing, and thorough documentation.
Apply programming skills in Python, PySpark, Scala, and SQL to effectively manipulate and transform data.
Understand and utilize cloud computing platforms and services offered by providers such as AWS, Azure, and Google Cloud.
Develop data pipelines using orchestration frameworks (e.g., Apache Airflow, Luigi, Mage, Databricks Workflows), integrating with PySpark and/or Scala as needed.
Understand and apply software design principles to data engineering projects.
Apply data modeling techniques to design efficient and scalable data models.
Optimize PySpark and SQL query performance, considering data volume, query complexity, and system resources.
Use data tools for ingestion, transformation, analysis, and visualization.
Soft Skills:
Logical reasoning and analytical skills;
Meeting deadlines and delivering quality work;
Effective and transparent communication with other teams and coworkers;
Communication and interpersonal skills to present and discuss cases with other involved areas;
Autonomy in performing tasks;