Deskripsi pekerjaan Data Engineer PT. Ebdesk Teknologi
-We leverage big data and Artificial Intelligence to deliver strategic insights for businesses and government institutions. As a Data Engineer, you will build scalable data infrastructure that powers our AI products and analytics platforms.
You will work closely with AI Engineers to develop reliable data pipelines, optimize data processing, and ensure high-quality datasets for AI model development and business intelligence.
Key Responsibilities
- Design, build, and maintain scalable ETL/ELT pipelines for structured and unstructured data.
- Develop automated data ingestion, scraping, transformation, and validation processes.
- Build and optimize relational and non-relational databases to support high-volume data processing.
- Collaborate with AI and Data Science teams to prepare, manage, and deliver datasets for machine learning and LLM applications.
- Develop robust data integration services using Python and SQL.
- Implement data quality monitoring, logging, and error handling across data pipelines.
- Optimize data workflows for performance, scalability, and reliability.
- Integrate multiple data sources through APIs, web services, and browser automation tools.
- Utilize Docker, Git, and Linux environments for deployment, automation, and version control.
- Contribute to improving the company's AI data platform and data engineering best practices.
Requirements
- Bachelor's degree in Computer Science, Information Technology, Data Science, or a related field.
- Strong understanding of relational and non-relational database concepts, including ERD design.
- Proficient in SQL (SELECT, INSERT, UPDATE, DELETE, JOIN, Aggregation, Window Functions).
- Strong proficiency in Python for data processing and automation.
- Experience building ETL/ELT pipelines and large-scale data processing workflows.
- Experience with web scraping tools such as BeautifulSoup, Selenium, or Playwright.
- Familiar with Python libraries including requests, asyncio, re, and pandas.
- Experience consuming REST APIs and handling HTTP protocols.
- Familiar with Linux environments, Docker, and Git version control.
- Understanding of asynchronous programming and message queue concepts.
- Experience implementing logging, monitoring, and debugging for production pipelines.
- Strong analytical thinking, problem-solving skills, and attention to detail.
Nice to Have
- Experience working with AI, Machine Learning, or Large Language Model (LLM) projects.
- Familiarity with vector databases, embedding pipelines, or Retrieval-Augmented Generation (RAG).
- Knowledge of workflow orchestration tools such as Airflow or Prefect.
- Experience with cloud platforms (AWS, GCP, or Azure).
- Understanding of MLOps or data platform architecture.
- Familiarity with Apache Kafka, RabbitMQ, or similar messaging systems.
Why Join Us?
- Work on impactful AI and Big Data solutions used by enterprise and government clients.
- Collaborate with experienced AI Researchers, Data Scientists, and Machine Learning Engineers.
- Opportunity to build scalable data platforms that power real-world AI applications.
- Fast-paced environment with exposure to cutting-edge technologies and continuous learning.

