An Agentic AI-Powered Linkedin Automation to fetch post and post them into Linkedin
An intelligent, fully-automated multi-stage pipeline that scrapes fresh job listings from LinkedIn & Indeed, ranks the best opportunities using Google Gemini AI, generates beautiful visual job-card images, stores everything in Supabase, and auto-posts to LinkedIn β all on a schedule via GitHub Actions.
- πΈ Project Overview
- β¨ Key Features
- βοΈ Pipeline Flow β How the System Works
- πΌοΈ Automation Pipeline
- π Folder Structure
- π Installation & Setup Guide
- π Environment Variables
- π€ GitHub Actions β Automation Workflows
- π οΈ Tech Stack
LinkedIn Job Automation is a production-grade, fully-automated job discovery and social media posting pipeline built entirely in Python. It is designed to run hands-free on a schedule, discovering the most relevant junior/mid-level tech job listings in Pakistan every day, enriching them with AI, generating branded visual cards, and posting them to LinkedIn automatically β without any manual effort.
| Stage | What Happens |
|---|---|
| π Scrape | Scrapes LinkedIn & Indeed daily for 5 job categories |
| π€ Rank | Google Gemini AI picks the single best job per category |
| π Enrich | Fetches full job descriptions, logos, salary data |
| βοΈ Summarize | Gemini AI distills each job into a clean LinkedIn-ready post |
| π¨ Design | Playwright renders beautiful color-coded job card images |
| βοΈ Store | Uploads images and all data to Supabase cloud |
| π€ Post | Buffer GraphQL API publishes posts to LinkedIn on schedule |
Target Audience: Job seekers, recruiters, and tech communities in Pakistan looking for curated, daily-fresh junior and mid-level tech opportunities.
Scrapes LinkedIn and Indeed simultaneously for 5 tech job categories (Full Stack, AI/Data, Mobile, UI/UX Design, Software/DevOps). Uses smart deduplication β enforcing 1 job per company per category and filtering out senior/director/manager roles automatically.
Uses Google Gemini gemini-3.5-flash-lite to evaluate every scraped job on company reputation, hiring prestige, market performance, and role seniority fit β then picks the single absolute best job per category.
Re-scrapes LinkedIn with linkedin_fetch_description=True to pull full job descriptions, company logos, company URLs, and salary/pay information for each ranked job.
Gemini AI processes each enriched job into a Pydantic-validated structured summary, extracting: job summary, key requirements, required skills, company perks, workplace type (Remote/Hybrid/On-site), smart hashtags, and a matching brand color code.
Uses Playwright (headless Chrome/Edge) to render a branded HTML/CSS job card template dynamically injected with each job's data and color scheme β then screenshots it as a PNG image, ready for social posting.
Uploads generated PNG images to Supabase Storage (job-images bucket) and inserts all structured job data into a Supabase PostgreSQL jobs table with pending status β creating a posting queue.
The LinkedIn posting pipeline fetches the oldest pending job from Supabase, downloads its image, formats a rich LinkedIn post text with emojis and hashtags, then publishes via the Buffer GraphQL API β and marks the job posted when complete.
Scrapers are known to be blocked by anti-bot filters in CI/CD environments. The pipeline includes a smart fallback cache system β if a scraper is blocked, pre-defined sample job data is injected so the AI ranking and downstream steps always complete successfully.
Three independent GitHub Actions workflows handle everything:
run-job.ymlβ Runs the full scrape β rank β enrich β summarize β design β store pipelinelinkdin-post.ymlβ Fetches pending jobs and posts to LinkedInremove-job.ymlβ Cleans up old/expired jobs
The system operates as two independent pipelines triggered by GitHub Actions (manually or via cron-job.org):
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β MASTER PIPELINE (main.py) β
βββββββββββββββββββββββββ¬ββββββββββββββββββββββββββββββββββββββββββ
β
βββββββββββββββββΌββββββββββββββββββ
β STEP 1 β Core Pipeline β
β Core/main-core.py β
ββββ¬βββββββββββββββββββββββββββ¬ββββ
β β
ββββββββββββΌβββββββββββ βββββββββββββΌβββββββββββ
β LinkedIn Engine β β Indeed Engine β
β linkdin-engine.py β β indeed-engine.py β
β β β β
β Scrapes 5 categories β β Appends 5 unique jobs β
β (15 raw β 5 unique) β β per category below β
β β Data/Job-Result- β β LinkedIn results β
β cache/*.json β β β
ββββββββββββββββββββββββ βββββββββββββββββββββββββ
β β
ββββββββββββ¬ββββββββββββββββ
β
βΌ [Fallback cache injected if scrapers blocked]
βββββββββββββββββββββββββββββββ
β Core/job-rank.py β
β Gemini AI ranks each β
β category β picks best job β
β β Data/Job_Rank.json β
βββββββββββββββ¬ββββββββββββββββ
β
βββββββββββββββΌββββββββββββββββ
β Core/job-info.py β
β Re-scrapes LinkedIn for β
β full description + logo β
β + pay info β
β β Data/Job-Info.json β
βββββββββββββββ¬ββββββββββββββββ
β
βββββββββββββββΌββββββββββββββββ
β Core/job-summery.py β
β Gemini AI β Pydantic schema β
β Structured LinkedIn post β
β data + color codes β
β β Data/Job-summery.json β
βββββββββββββββ¬ββββββββββββββββ
β
βββββββββββββββΌββββββββββββββββ
β STEP 2 β Design Generator β
β Job-Post-Design/ β
β main-job-post.py β
β β
β post-data.py β batches β
β path-finder.py β Playwright β
β renders HTML card β PNG β
β β Generated-Images/*.png β
βββββββββββββββ¬ββββββββββββββββ
β
βββββββββββββββΌββββββββββββββββ
β STEP 3 β Storage β
β Main-Storage/main-storage.pyβ
β Merges image paths + data β
β β main-storage.json β
βββββββββββββββ¬ββββββββββββββββ
β
βββββββββββββββΌββββββββββββββββ
β STEP 4 β Supabase Sync β
β Main-Storage/Supabse.py β
β Uploads PNG to Storage β
β Inserts job record to DB β
β status = "pending" β
βββββββββββββββββββββββββββββββ
ββββββββββββββββββββββββββββββββββββββββββββββββββββ
β LINKEDIN POSTING PIPELINE β
βββββββββββββββββββββββββ¬βββββββββββββββββββββββββββ
β
βββββββββββββββββΌββββββββββββββββββ
β post.py β
β Fetches oldest "pending" job β
β from Supabase β downloads image β
β β saves post.json locally β
βββββββββββββββββ¬ββββββββββββββββββ
β
βββββββββββββββββΌββββββββββββββββββ
β linkdin-post.py β
β Formats rich post text β
β with emojis + hashtags β
β β Buffer GraphQL API β
β β Publishes to LinkedIn β
β β Updates Supabase β
β status = "posted" β
βββββββββββββββββββββββββββββββββββ
| # | Category | Output Cache File |
|---|---|---|
| 1 | Frontend, Full Stack & Backend | Fullstack.json |
| 2 | AI & Data Engineering | AIEngineer.json |
| 3 | Mobile Developer | MobileDeveloper.json |
| 4 | UI/UX & Graphic Designer | Designer.json |
| 5 | Software & DevOps Engineering | SoftwareEngineer.json |
Each job category gets a distinct brand color injected into the visual card:
| Category | Color Name | Hex Code |
|---|---|---|
| Full Stack | Deep Teal | #1A6B72 |
| AI Engineering | Muted Violet | #5B4A8A |
| Design | Steel Blue | #2E6B8A |
| Software Engineering | Forest Green | #3D6B52 |
| Other / Default | Dusty Plum | #6B4E71 |
Linkdin-Automation/
β
βββ main.py # Master entry point β runs the full pipeline
βββ buffer.py # Buffer API utility (channel/org ID helper)
βββ requirements.txt # All Python dependencies
βββ .env # Local secrets (NEVER commit this)
βββ .gitignore # Git ignore rules
β
βββ Core/ # Core job discovery & AI processing pipeline
β βββ main-core.py # Orchestrates all core steps with resilience
β βββ job-rank.py # Gemini AI ranks jobs, picks best per category
β βββ job-info.py # Re-scrapes LinkedIn for full job enrichment
β βββ job-summery.py # Gemini AI structured Pydantic summarization
β β
β βββ Job-search/ # Web scraper engines
β βββ linkdin-engine.py # LinkedIn scraper (5 categories, deduped)
β βββ indeed-engine.py # Indeed scraper (appends to LinkedIn cache)
β βββ job_quries.py # Search query definitions & role keywords
β
βββ Data/ # Pipeline data cache (auto-generated)
β βββ Job-Result-cache/ # Raw scraper output (per-category JSON files)
β β βββ Fullstack.json
β β βββ AIEngineer.json
β β βββ MobileDeveloper.json
β β βββ Designer.json
β β βββ SoftwareEngineer.json
β βββ Job_Rank.json # Gemini-ranked best job per category
β βββ Job-Info.json # Enriched job details (description, logo, pay)
β βββ Job-summery.json # Final structured summaries ready for posting
β
βββ Job-Post-Design/ # Visual job card image generator
β βββ main-job-post.py # Orchestrates design pipeline steps
β βββ post-data.py # Batches job data into Post-data.json
β βββ path-finder.py # Playwright renderer β PNG screenshots
β βββ Post-data.json # Batched post data (auto-generated)
β β
β βββ Design-Template/
β β βββ index.html # HTML/CSS job card template (dynamic)
β β
β βββ Generated-Images/ # Output PNG job card images (auto-generated)
β βββ *.png
β
βββ Main-Storage/ # Storage & Supabase sync layer
β βββ main-storage.py # Merges data + image paths β main-storage.json
β βββ Supabse.py # Uploads images + inserts jobs to Supabase
β βββ main-storage.json # Final merged dataset (auto-generated)
β
βββ Linkdin-posting/ # LinkedIn auto-posting pipeline
β βββ main-post.py # Orchestrates the posting workflow
β βββ post.py # Fetches pending job from Supabase queue
β βββ linkdin-post.py # Formats text + publishes via Buffer GraphQL
β β
β βββ Post-Data/ # Temporary posting data (auto-generated)
β βββ post.json
β βββ post_image.png
β
βββ Data-Remove/ # Cleanup utilities
β βββ remove.py # Removes old/expired jobs from Supabase
β
βββ .github/
βββ workflows/ # GitHub Actions automation
βββ run-job.yml # Triggers the full pipeline
βββ linkdin-post.yml # Triggers LinkedIn posting
βββ remove-job.yml # Triggers job cleanup
- Python 3.10 β 3.12 β Download here
- Git β Download here
- Google Chrome or Microsoft Edge (required for Playwright screenshots)
https://github.com/Developer359/Linkdin-Automation.git
cd Linkdin-Automationpip install python-jobspy pandas python-dotenvpip install google-genaipip install playwright
playwright installpip install supabase python-dotenvOr install everything at once from requirements.txt:
pip install -r requirements.txt
playwright installCreate a .env file in the root of the project and fill in your credentials:
# Google Gemini AI
GEMINI_API_KEY=your_gemini_api_key_here
# Supabase β Database & Storage
SUPABASE_URL=https://your-project.supabase.co
SUPABASE_KEY=your_supabase_service_role_key_here
# Buffer β LinkedIn Auto-Posting
BUFFER_API_KEY=your_buffer_api_key_here
BUFFER_CHANNEL_ID=your_buffer_linkedin_channel_id_here| Key | Where to Get It |
|---|---|
GEMINI_API_KEY |
Google AI Studio β Create API Key |
SUPABASE_URL |
Supabase Dashboard β Settings β API β Project URL |
SUPABASE_KEY |
Supabase Dashboard β Settings β API β service_role secret key |
BUFFER_API_KEY |
Buffer Developer Portal β Access Token |
BUFFER_CHANNEL_ID |
Run python buffer.py β prints your connected channel IDs |
In your Supabase project, go to SQL Editor and run:
CREATE TABLE jobs (
id BIGSERIAL PRIMARY KEY,
created_at TIMESTAMPTZ DEFAULT NOW(),
job_title TEXT,
company TEXT,
location TEXT,
job_url TEXT,
company_email TEXT,
source TEXT,
pay_info TEXT,
job_summary TEXT,
requirements JSONB,
required_skills JSONB,
what_we_offer JSONB,
about_company TEXT,
workplace_type TEXT,
image_url TEXT,
status TEXT DEFAULT 'pending'
);Then in Supabase β Storage, create a Public bucket named job-images.
python main.pypython Linkdin-posting/main-post.py| Variable | Required | Description |
|---|---|---|
GEMINI_API_KEY |
β Yes | Google Gemini AI API key (raw key only, no prefix) |
SUPABASE_URL |
β Yes | Your Supabase project URL |
SUPABASE_KEY |
β Yes | Supabase service role key (has full DB access) |
BUFFER_API_KEY |
β Yes | Buffer access token for GraphQL API |
BUFFER_CHANNEL_ID |
β Yes | Buffer LinkedIn channel ID for posting |
β οΈ Security Warning: Never commit your.envfile to Git. It is already listed in.gitignore. For GitHub Actions, add all secrets via: Repository β Settings β Secrets and Variables β Actions β New repository secret.
Runs python main.py β the complete scrape β rank β enrich β design β store pipeline.
Required Secrets: GEMINI_API_KEY, SUPABASE_URL, SUPABASE_KEY
Trigger: Manual (workflow_dispatch) or via cron-job.org API call for daily scheduling.
Runs python Linkdin-posting/main-post.py β fetches the next pending job and posts to LinkedIn.
Required Secrets: SUPABASE_URL, SUPABASE_KEY, BUFFER_API_KEY, BUFFER_CHANNEL_ID
Trigger: Manual (workflow_dispatch) or scheduled via cron-job.org.
Runs python Data-Remove/remove.py β wipes all job images from Supabase Storage and clears the entire jobs database table, keeping your storage clean and the ID counter reset before the next pipeline run.
What it does, step by step:
- ποΈ Clears the
job-imagesStorage bucket β lists every file in the bucket and bulk-deletes them all (skips the hidden.emptyFolderPlaceholderfile automatically) - ποΈ Truncates the
jobsdatabase table β calls a Supabase PostgreSQL RPC functionreset_jobs_tablethat truncates the table and resets theidauto-increment sequence back to1 - β¨ Leaves Supabase completely clean β ready for the next fresh pipeline run
Required Secrets: SUPABASE_URL, SUPABASE_KEY, SUPABASE_SERVICE_ROLE_KEY
Trigger: Manual (workflow_dispatch) or scheduled via cron-job.org before each new pipeline run.
β οΈ Important: Before running this workflow, you must create thereset_jobs_tableRPC function in your Supabase SQL Editor:CREATE OR REPLACE FUNCTION reset_jobs_table() RETURNS void AS $$ BEGIN TRUNCATE TABLE jobs RESTART IDENTITY; END; $$ LANGUAGE plpgsql;
- Go to your repository on GitHub
- Click Settings β Secrets and variables β Actions
- Click New repository secret and add each of the following:
GEMINI_API_KEY β Your raw Google AI Studio key (starts with AIza...)
SUPABASE_URL β https://your-project.supabase.co
SUPABASE_KEY β Your service_role key from Supabase
BUFFER_API_KEY β Your Buffer access token
BUFFER_CHANNEL_ID β Your Buffer LinkedIn channel ID
| Technology | Role |
|---|---|
| Python 3.10 β 3.12 | Core language for all pipeline logic |
Google Gemini AI (gemini-3.5-flash-lite) |
Job ranking, structured summarization, NLP |
| python-jobspy | Multi-site job scraping (LinkedIn + Indeed) |
| Playwright | Headless browser for HTMLβPNG job card generation |
| Supabase | PostgreSQL database + S3-compatible image storage |
| Buffer GraphQL API | LinkedIn post publishing & scheduling |
| Pydantic | Strict schema validation for Gemini AI outputs |
| GitHub Actions | CI/CD orchestration & cloud automation |
| python-dotenv | Secure environment variable management |
| pandas | DataFrame processing for scraper results |