Job Description
Data Development Engineer (Real-time Computing Focus)
Position Overview:
- The company has established a well-developed data warehouse and real-time data system with capabilities in data collection, computation, storage, and services. With the growth of business scale and the increase in real-time application scenarios, we are looking to hire two mid-to-senior data development engineers to take charge of real-time data demand delivery, task optimization, and engineering governance.
- This role primarily focuses on real-time data development while also contributing to offline data warehouse construction. You will address critical issues in real-time data pipelines, such as data latency, state expansion, resource consumption, data quality, and consistency, driving the real-time data system toward greater stability, efficiency, and standardization.
Key Responsibilities
- Analyze, design, and develop real-time data solutions using technologies like Flink and Kafka to support real-time analytics, monitoring, metrics, and business decision-making.
- Participate in real-time data warehouse construction and model design, ensuring data layering, metric definitions, dimensional modeling, and data services align with business needs.
- Continuously optimize existing real-time tasks to resolve issues like data latency, backpressure, Checkpoint anomalies, state expansion, resource inefficiency, and data inconsistency, improving task stability and computational efficiency.
- Engage in data governance and engineering quality initiatives, including task dependency mapping, metadata management, data quality monitoring, alert mechanisms, development standards, and troubleshooting frameworks.
- Collaborate with upstream and downstream teams to maintain the stability, accuracy, and timeliness of data pipelines, while also assisting in the daily maintenance of big data platforms and computing engines.
- Document and share real-time development best practices and technical insights through code reviews, technical discussions, and solution evaluations to enhance the team's real-time development capabilities.
Job Requirements
- Bachelor’s degree or higher in Computer Science, Software Engineering, Information Technology, or a related field, with 5+ years of experience in data development or big data engineering (exceptions may be made for highly skilled candidates).
- Proven experience in data warehouse projects, with expertise in real-time data warehouse architecture, data layering, and model design to translate business requirements into technical solutions.
- Proficiency in Flink, including hands-on experience with Flink SQL and DataStream API, and a deep understanding of State, Checkpoint, Watermark, windowing, backpressure, and fault tolerance mechanisms.
- Experience in real-time task development and optimization, with the ability to independently troubleshoot issues like latency, high resource consumption, state anomalies, data skew, and consistency.
- Familiarity with Kafka and real-time data pipelines, covering the full lifecycle of data collection, transmission, computation, storage, and service delivery.
- Knowledge of Hadoop, Hive, Spark, and other big data technologies, with some experience in offline data development.
- Strong SQL skills for complex data processing and performance tuning, along with proficiency in at least one programming language (Python, Java, or Scala).
- Experience with databases or storage systems like MySQL, PostgreSQL, or Elasticsearch.
- Excellent business acumen, communication skills, and written expression, with a strong sense of responsibility and initiative to drive problem-solving and project execution.
Preferred Qualifications
- Experience in large-scale Flink task development, performance tuning, or cluster governance.
- Background in real-time data warehouse optimization, pipeline governance, or technical upgrades.
- Familiarity with real-time OLAP technologies like Hologres, ClickHouse, or Doris.
- Exposure to complex business scenarios such as real-time metrics, monitoring, risk control, or recommendation systems.
Our Ideal Candidate
We seek someone who not only delivers data solutions but also delves into real-time computing mechanisms, prioritizing pipeline stability, performance, and engineering quality. When tackling complex issues, you should be able to diagnose root causes from multiple angles—data, code, task configurations, and operational mechanisms—and formalize solutions into reusable standards or tools.
Benefits
Details to be discussed during the interview process.