Описание
Technology->Analytics - Packages->Python - Big Data,Technology->Big Data - Data Processing->PySpark Design, develop, and maintain scalable batch/stream data pipelines using Python and PySpark in distributed environments. Implement efficient transformations, aggregations, and joins on large datasets while ensuring performance and cost optimization. Write optimized SQL for data extraction, validation, and reconciliation across multiple sources.
Build reusable, testable modules and follow engineer…
Контакты работодателя (email/phone/telegram) скрыты из публичного превью —
отправьте резюме, чтобы мы связали вас напрямую.