Authorization enforced in the frontend only
Row-level security policies were permissive; tenant isolation existed in the UI, not in the database. Any authenticated user could have queried another client's documents directly against the API.
The client's product team had done something that would have been impossible two years ago: they built a working AI compliance system themselves.
Using Lovable, their subject-matter experts, compliance consultants, not developers, turned years of regulatory know-how into a functioning application. Document upload, automated policy checks against current regulatory frameworks, an LLM-driven gap analysis, and a reporting layer their consultants could actually use. It demoed beautifully. Internal stakeholders were convinced. A pilot customer was waiting.
Then came the questions their prototype couldn't answer. Where is the client data stored? What happens when two consultants edit the same assessment? Who can see which tenant's documents? What does the system do when the AI returns something unexpected?
These aren't questions a prototype is built to answer. They are exactly the questions a compliance customer asks first, and this client's entire business rests on being credible about them.
They approached WOLKK to close the gap between a convincing demo and a system they could sell.
We started with a one-week audit before writing a single line of production code. The prototype was better than most, the domain logic was thoughtful, the UX genuinely reflected how consultants work. But the findings were typical of AI-generated codebases at this stage:
Row-level security policies were permissive; tenant isolation existed in the UI, not in the database. Any authenticated user could have queried another client's documents directly against the API.
Malformed LLM responses, oversized uploads and rate limits all produced a blank screen rather than a recoverable error.
The same regulatory ruleset implemented three times with three slightly different behaviours.
We delivered the audit as a prioritised written report: what was a launch blocker, what was technical debt, and what was fine as it stood. Roughly half the codebase, in our assessment, needed no rework at all.
The client's domain logic was their competitive advantage. We preserved it, consolidated the duplicated implementations into a single tested rules engine, and documented the regulatory mappings so they could be maintained by the compliance team rather than only by engineers.
Server-side authorization, database-level row security with enforced tenant isolation, secrets moved into a managed vault, all LLM calls routed through a backend service. We then had the tenant isolation independently penetration-tested before go-live.
EU-region hosting, an EU-resident model endpoint under a DPA, defined retention and deletion paths for uploaded documents, complete audit logging, and a sub-processor register their own compliance team could hand to a customer without caveats.
Structured output with schema validation, retry and fallback behaviour, an evaluation suite that runs known documents against known expected findings so a prompt change can't quietly degrade assessment quality.
Unit and integration tests over the rules engine and authorization layer, CI on every pull request, staging environment, monitored production deployment with rollback.
Architecture documentation, decision records, runbooks, and a working session with the client's newly hired in-house developer, who took over day-to-day maintenance with the codebase in a state he could actually read.
The system went live with the pilot customer and has since been rolled out across the firm's client base. The client's product team continues to prototype new modules in Lovable; WOLKK takes them through the same hardening path before they reach production. That has become the working model: their experts move fast on the idea, our engineers make it hold.
AI development environments have moved the bottleneck. Getting to a working prototype is no longer the hard part, and domain experts building their own tools is a genuinely good development.
But the last stretch to production is still engineering work: authorization, data protection, failure handling, testability, maintainability. It isn't glamorous and AI tools don't do it well, because the person prompting usually doesn't know to ask for it.
That stretch is what we do.
Have a prototype that needs to become a product? Start with an audit. We'll tell you what's solid, what's a risk, and what it actually takes to ship it.