An LLM-powered software engineering framework that transforms
natural-language requirements into fully functional Django applications — routed through
a Structured Intermediate Representation (SIR) and scored by an automated,
multi-dimension quality evaluator before it's handed back.
DjangoForge takes a plain-language app description and turns it into a real, runnable
Django project. The pipeline is labeled in its own architecture as
NL → SIR → LLM Synthesis → Automated Evaluation → Live Preview / ZIP —
natural language in, a structured intermediate representation in between, multi-file
Django code out, automatically scored before it's returned.
Problem
Scaffolding a new Django app — models, views, URLs, admin registration, migrations,
templates — is repetitive groundwork that has to happen before any real feature work can
start. Generating that code with an LLM directly from a raw prompt is also unreliable on
its own: without structure in between, output quality and consistency vary a lot.
DjangoForge targets both problems at once.
Solution
A prompt box takes a description — typed freely, in English or Arabic, or picked from
example prompts like "a task manager with projects and tasks" — and the backend converts it
into a structured JSON schema (the SIR) before generating code. The result: full file
structure, models, views, URLs, templates, migrations, and seed data, with a live in-browser
preview and a one-click .zip download.
Prompt interface with example descriptions and file explorer
Key Features
NL-to-SIR extraction: converts an unstructured prompt into a clean
JSON schema defining models, fields, view types, and dynamic UI theme.
Multilingual input: free-text prompts including non-English input —
the demo shows a full app generated from an Arabic description.
Automated quality evaluation: every generation is scored across five
weighted dimensions before being handed back — see Engineering Highlights.
Real-time SSE generation: built on FastAPI with Server-Sent Events
for low-latency streaming instead of a single blocking request.
File explorer: generated projects appear as a real file tree —
manage.py, requirements.txt, app folder with
models.py, views.py, urls.py, admin.py,
migrations, and seed data.
Code & Eval tabs: inspect generated source directly, alongside the
evaluation score breakdown for that generation.
Download: export the generated project as a .zip.
Example: an Arabic prompt generating a 19-file "flower agenda" Django app, with live preview and a 97.3% OQS score
Natural Language Input — the user's free-text description.
SIR Extraction Engine — converts the prompt into a Structured
Intermediate Representation (a normalized JSON schema of models, fields, and views).
LLM Multi-File Synthesizer — generates models, views, URLs,
templates, and a matching UI preview from the SIR.
Automated Evaluation Engine — computes an Overall Quality Score
(OQS) across five weighted dimensions.
Live Preview & ZIP Output — renders the generated app in-browser
and packages it for download.
The backend is a FastAPI service that calls the Claude API for generation, with
Server-Sent Events streaming the response back to a lightweight static frontend.
Structured Intermediate Representation: rather than generating code
directly from a raw prompt, the pipeline passes through a structured JSON schema first —
reducing inconsistency in the generated output and giving the synthesizer a
well-defined contract to generate against.
Five-dimension automated evaluation: every generated app is scored on
Syntax Correctness (30%, via AST parsing across all Python modules), Field & Model
Consistency (25%), View Coverage (20%), Template Completeness (15%), and URL Integrity
(10%) — a real, if partial, stand-in for the manual code review a generated app would
otherwise need.
Streamed generation: built on FastAPI with Server-Sent Events, so
multi-file generation streams back progressively instead of leaving the UI blocked on one
long request.
Rate-limited API surface: request throttling (SlowAPI) protects the
underlying LLM API from abuse on a publicly reachable endpoint.
Multilingual generation: the SIR extraction step accepts prompts in
languages other than English, as demonstrated with Arabic input.
Results
In the demonstrated example, a single Arabic-language prompt ("I want a pink app for
a flower agenda") produced a complete 19-file Django application — models, views, URLs,
admin config, and seed data — with a working live preview and an Overall Quality Score of
97.3% from the automated evaluator. No broader benchmark numbers (average
OQS across many prompts, generation latency at scale) are published here.