The Problem

Arabic newsrooms face a dual challenge:

  • Speed pressure – Breaking news demands fast content turnaround
  • Quality Control – AI-generated Arabic often lacks grammar, regional tone, and editorial standards.
  • Accountability – No audit trail when AI content goes wrong

Most AI writing tools give you one draft and call it done.
Editors either accept flawed output or rewrite from scratch.

My Solution

I built a 3-agent agentic pipeline where AI critiques and refines its own work before presenting options to the editor.

The Pipeline

Key Features

  • Two Generation Modes
    • Simple Mode
    • Agentic Mode
  • Arabic – First Design
    • Right-to-left interface
    • Arabic prompts and critique feedback
    • Designed for Gulf newsroom standards
  • Human-in-the-loop Governance
    • Editor always has final say
    • Override tracking: system logs when editors change AI output
    • Override percentage calculated and stored

Full Audit Trail

Every generation stores:

  • Original Input
  • Draft V1 (Generator output)
  • Critique (Critic feedback with scores)
  • Draft V2 (Refiner output)
  • Selected version (1 or 2)
  • Editor’s final body (if edited)
  • Override delta and percentage
  • Timing Metrics (gen_time_ms, critique_time_ms, refine_time_ms)

Archietecture

Evaluation Framework

AI Output is unpredictable. Every generation could fail silently. Editors catching mistakes is expensive.
I built a 20 case eval suite that validates LLM ouput against the same rules the production system enforces.

The Eval Set

CheckTest Cases
Headlines5
Tags5
Body Lenght5
Slug5
SEO description5

Each case is a real editorial – style Arabic topic like:

Structural Checks

RuleValidation
HeadlinesExactly 3
TagsExactly 5
Body Length350-650 words
SlugArabic-only with hyphens
SEO description<160 characters

Running Evals

node server/evals/run-evals.js http://localhost:5001

Output:

[PASS] #1 (Politics) score=8/10
[PASS] #2 (Politics) score=7/10
[FAIL] #3 (Politics)
        - body word count 312 outside [350, 650]
...
18/20 passed (90.0%)

Exit code 0 = all pass. Integrates with CI.

Why This Matters

Most AI writing tools ship without evaluation. When output fails, editors catch it – or worse, readers do.
This eval framework:

  • Catches structural failures before they reach the editor
  • Enables CI/CD – blocks deploys if evals regress
  • Documents Quality – 90% pass rate is auditable proof the system works

Tech Stack:

  • Frontend: Vanilla JS, Tailwind CSS
  • Backend: Node.js, Express
  • Database: SQLite
  • AI: Google Gemini 2.5 Flash (architected to swap to Jais Arabic LLM)

Why this Matters for Saudi Newsrooms:

ChallengeHow this Solves it
AI makes grammar errorsCritic agent catches issues before editor sees them
No visibility into AI decisionsFull audit trail for every generation
Editors waste time fixing AIV2 is already improved; editor reviews, not rewrites
Compliance concernsOverride tracking proves human oversight

Screenshots:

➡️ Mode Toggle

➡️ Agent Progress

➡️ Critique Panel

➡️ Side – by – Side Comparison

➡️ Demo Video

What I learned

  • Agentic loops need structure – Without clear handoffs between agents, output degrades
  • Arabic critique is harder – The critic prompt required careful tuning for proper Arabic Feedback
  • Audit trails build trust – Editors adopted faster when they could see exactly what the AI changed

Links


Built With

  • Node.js + Express
  • SQLite
  • Google Gemini 2.5 Flash
  • Vanilla JS + Tailwind CSS
  • Figma (architecture diagrams)

Part of myAI PM portfolio demonstrating agentic AI patterns for enterprise content workflows.