Software Engineer II - Backend & Data Platform (Python, Kafka, PySpark)

HiLabs
HiLabs

Software Engineering · Full-time

Pune, Maharashtra, India

Posted on Sep 4, 2026

HiLabs builds AI-driven data solutions for US healthcare payers, solving large-scale data quality problems across provider, claims and clinical datasets. Our engineering teams in Pune build the platform that powers this at production scale.

About the Role

We are hiring a Software Engineer II to build scalable, reliable, production-grade backend services for our core data platform. The work spans Python backend engineering, distributed data processing with PySpark, event-driven systems on Kafka and Redis, and ML inference integration — all within a secure AWS environment. You will work closely with Data Science, Data Engineering, DevOps/SecOps and Product teams.

What You Will Do

  • Design, build and maintain Python backend services and microservices using FastAPI/Flask.
  • Build scalable PySpark pipelines for ingestion, classification, linkage, feature engineering and schema generation.
  • Develop event-driven and asynchronous services using Kafka and Redis Streams — producers, consumers, consumer groups, topics, partitioning, offset management and retry handling.
  • Use Redis for caching, fast lookups, TTL-based state and distributed coordination.
  • Build and deploy containerized services on AWS EKS / Kubernetes.
  • Integrate with AWS S3, EMR, EventBridge, PostgreSQL, Snowflake, Kafka and Redis.
  • Deploy batch scoring and ML inference services using versioned model artifacts.
  • Implement reliability patterns: idempotency, retries, timeouts, dead-letter handling and failure recovery.
  • Design efficient database schemas, indexes, queries and data-access layers.
  • Implement logging, metrics, monitoring and alerting; write unit and integration tests; participate in code reviews.
  • Troubleshoot across application, data, messaging and infrastructure layers.

Must Have

  • 3–5 years of hands-on backend software engineering experience.
  • B.E./B.Tech/M.Tech/MCA in Computer Science or a related field from a Tier-1 institute (IIT, NIT, BITS, IIIT or equivalent). This is a mandatory requirement for this position.
  • Strong Python — OOP, data structures, design principles.
  • PySpark / Apache Spark and distributed data processing at scale.
  • Apache Kafka — producers, consumers, consumer groups, partitions and offset management.
  • Redis — caching, TTL, Pub/Sub or Streams.
  • FastAPI or Flask, REST APIs, microservices.
  • AWS — S3, IAM, CloudWatch, and EKS or EMR.
  • Docker, with working knowledge of Kubernetes.
  • PostgreSQL / SQL — schema design, indexing, query optimization.
  • Git, CI/CD, unit and integration testing.

Good to Have

  • Amazon MSK or Kafka on Kubernetes; Redis Cluster or ElastiCache.
  • AWS EMR and Spark at scale; Snowflake.
  • Parquet and schema management.
  • ML inference, model serving, embeddings, risk-scoring or MLflow.
  • GPU-based workloads.
  • Healthcare, claims or clinical data; PHI and HIPAA awareness.

Ideal Candidate

A strong backend and distributed-systems engineer comfortable across application, data and infrastructure layers, with real ownership of production services — debugging depth, system-design fundamentals, and the instinct to build fault-tolerant, observable systems.

Location: Pune, Kharadi (on-site). Candidates open to relocating to Pune may apply.