Skip to content
View akarce's full-sized avatar

Highlights

  • Pro

Block or report akarce

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
akarce/README.md

Ceyhun Akar

Data engineer in Istanbul. I build streaming pipelines and the secure, multi-tenant analytics layer on top of them.

Now: Data Engineer at Mediastream AG, working on ComplyGuard (compliance analytics for regulated iGaming). Kafka and PyFlink streams, Airflow and Spark on Kubernetes, ClickHouse with row-level security, FastAPI and SSO. That code is private, so I describe the work on my site.
Studying: Executive MBA at Istanbul Technical University.
Links: akarce.github.io · CV (PDF) · LinkedIn · Medium · YouTube

Open-source projects (2024)

Project What it does Stack
real-time-data-pipeline-kafka-mongo-elasticsearch-pyspark Streams Yelp reviews through Kafka and Spark, labels sentiment with DistilBERT, shows it in Kibana. 82-min build video Kafka, Spark, MongoDB, Elasticsearch
e2e-structured-streaming Airflow to Kafka to Spark Structured Streaming to Cassandra, all in Docker Compose. Article Airflow, Kafka, Spark, Cassandra
elk-stack-mastery Multi-node Elasticsearch, Logstash and Kibana with hot, cold and frozen lifecycle tiers. Video Elasticsearch, Logstash, Kibana
RedditDataPipeline Reddit API into MinIO and Hive through NiFi and Airflow, queried with Trino, shown in Superset. Article NiFi, Airflow, Hive, Trino

Latest writing

Latest videos

Popular repositories Loading

  1. e2e-structured-streaming e2e-structured-streaming Public

    End-to-end data pipeline that ingests, processes, and stores data. It uses Apache Airflow to schedule scripts that fetch data from an API, sends the data to Kafka, and processes it with Spark befor…

    Python 23 7

  2. elk-stack-mastery elk-stack-mastery Public

    A comprehensive project focusing on setting up and configuring the Elastic Stack (Elasticsearch, Logstash, and Kibana) for efficient log management and analytics. This project includes Elasticsearc…

    10

  3. Udacity-Data-Pipeline-with-Airflow Udacity-Data-Pipeline-with-Airflow Public

    Udacity Data Engineering Nanodegree Program, Data Pipeline with Airflow project using MinIO and Postgresql.

    Python 5 2

  4. RedditDataPipeline RedditDataPipeline Public

    Data Engineering with Reddit Api, Airflow, Hive, Postgres, MinIO, Nifi, Trino, Tableau and Superset

    Python 5 2

  5. real-time-data-pipeline-kafka-mongo-elasticsearch-pyspark real-time-data-pipeline-kafka-mongo-elasticsearch-pyspark Public

    A real-time data pipeline project using Kafka, MongoDB, Elasticsearch, and PySpark. Streams raw data from Kafka, enriches it with sentiment analysis using Hugging Face models, stores results in Mon…

    Jupyter Notebook 3 1

  6. e2e-otp-pipeline e2e-otp-pipeline Public

    End to End OTP Pipeline Project using Docker, Airflow, Kafka, KafkaUI, Cassandra, MongoDB, EmailOperator, SlackWebhookOperator and DiscordWebhookOperator

    Python 2 3