We are seeking a skilled PySpark Data Engineer to join our team and drive the development of robust data processing and transformation solutions within our data platform. You will be responsible for designing, implementing, and maintaining PySpark-based applications to handle complex data processing tasks, ensure data quality, and integrate with diverse data sources. The ideal candidate possesses strong PySpark development skills, experience with big data technologies, and the ability to work in a fast-paced, data-driven environment.

Key Responsibilities:Data Engineering Development:

Design, develop, and test PySpark-based applications to process, transform, and analyze large-scale datasets from various sources, including relational databases, NoSQL databases, batch files, and real-time data streams.
Implement efficient data transformation and aggregation using PySpark and relevant big data frameworks.
Develop robust error handling and exception management mechanisms to ensure data integrity and system resilience within Spark jobs.
Optimize PySpark jobs for performance, including partitioning, caching, and tuning of Spark configurations.

Data Analysis and Transformation:

Collaborate with data analysts, data scientists, and data architects to understand data processing requirements and deliver high-quality data solutions.
Analyze and interpret data structures, formats, and relationships to implement effective data transformations using PySpark.
Work with distributed datasets in Spark, ensuring optimal performance for large-scale data processing and analytics.

Data Integration and ETL:

Design and implement ETL (Extract, Transform, Load) processes to ingest and integrate data from various sources, ensuring consistency, accuracy, and performance.
Integrate PySpark applications with data sources such as SQL databases, NoSQL databases, data lakes, and streaming platforms

Qualifications and Skills:

Bachelors degree
in Computer Science, Information Technology, or a related field.
5+ years
of hands-on experience in big data development, preferably with exposure to data-intensive applications.
Strong understanding of
data processing principles
, techniques, and best practices in a big data environment.
Proficiency in PySpark, Apache Spark, and related big data technologies
for data processing, analysis, and integration.
Experience with
ETL development
and data pipeline orchestration tools (e.g., Apache Airflow, Luigi).
Strong analytical and problem-solving skills, with the ability to translate business requirements into technical solutions.
Excellent communication and collaboration skills to work effectively with data analysts, data architects, and other team members.

More Jobs at PHOTON

Technical Lead_Offshore

Hyderabad, Telangana, India

10.0 - 12.0 yrs

INR 3 - 10 Lacs

DotNet Developer

Chennai, Tamil Nadu, India

3.0 - 6.0 yrs

INR 3 - 6 Lacs

TFS Developer

Hyderabad, Telangana, India

2.0 - 4.0 yrs

INR 3 - 8 Lacs

Sr Data Architect

Chennai, Tamil Nadu, India

8.0 - 13.0 yrs

INR 8 - 13 Lacs

Data Engineer

Hyderabad, Telangana, India

3.0 - 7.0 yrs

INR 3 - 7 Lacs

Mock Interview

Practice Video Interview with JobPe AI

Start Job-Specific Interview

Start Your Job Search Today

Browse through a variety of job opportunities tailored to your skills and preferences. Filter by location, experience, salary, and more to find your perfect fit.

Job Application AI Bot

Apply to 20+ Portals in one click

Download Now

Download the Mobile App

Instantly access job listings, apply easily, and track applications.

Enhance Your Skills

Practice coding challenges to boost your skills

Start Practicing Now

PHOTON

Login to

Please Verify Your Phone or Email

Confirm Action

Search

Profile

Upskill and Grow with AI

SPARK Data Onboarding Engineer