Files
krow_backend/scripts/gen_resources.py
Suriyakumarvijayanayagam dc785b917c
Some checks failed
CI / test (push) Failing after 4m41s
CI / fixture (push) Failing after 8s
Separate what a worker does from what a company needs filled
Owliver could offer neither create. The Create Position flow worked and no chip
anywhere suggested it, because the chip row is entirely the backend's static
catalogue and no intent in it wrote anything. The gap was never in the
frontend's trigger matching — every phrasing already routed.

`employee_roles` is the supply side of `job_postings`. A posting is what the
ORGANIZATION needs filled; this is what a WORKER says they do. They share a
vocabulary and almost nothing else: "3 years" on a posting is a minimum an
applicant must clear, and the same words here are what the person has. There is
deliberately no foreign key between them — supply and demand already meet
through `job_applications`, which carries the funnel, the interview and the
outcome, and a second weaker link would disagree with it the first time
somebody withdrew.

NO NEW COMPANY ENTITY, AND THAT IS THE LOAD-BEARING DECISION. "Create a company
position" reads like it needs a client record. `organizations` is the TENANT —
absent from the resource table, absent from the policy map, written only by the
seeder — so creating a row there from a chat flow would provision a new tenant,
and the position would carry an org_id the operator's session cannot see. The
operator could never view the record they just created. That breaks I5 and I1
to add a feature nobody asked for. The client stays free text on the posting,
per blueprint decision D2, and the flow simply offers the clients this
organization already staffs for as chips. No schema change, no endpoint change.

Create is operators-only, and that is an I1 decision rather than a deferral.
The worker is named explicitly on the row and is deliberately NOT derived from
the session, because an operator recording a role on somebody's behalf is the
whole point of the flow. Granting talent the same Create would let a talent
caller write a role under any worker_email in the tenant — the attribution hole
Phase 3D closed elsewhere. Talent reads its own via a ScopeEmail predicate,
which is in place now so the grant is one line when a talent console exists.

`created_by` is in gen_resources.py's SERVER_OWNED as well as the policy's
Derived list. Both are required and the pairing is easy to miss: Derived fills
the column from the session, SERVER_OWNED is what makes the descriptor ReadOnly
so a request body cannot set it in the first place. Without it,
TestDerivedColumnsAreReadOnlyOrTalentScoped fails — verified by mutation, not
by reading.

The two catalogue intents carry PHRASE terms only. A bare "position" or "role"
term scores 10, the same as every reading on that page, and wins the tie on
declaration order — so a create chip would have arrived by evicting
`positions-attention` from the exact ordered result TestPositionsSuggestions
asserts. An offer to create something must not displace the reading a person
actually asked for. Neither declares a Subject, on the precedent of
`position-spec-steps`: a Subject would let the bare query "summarize" match
through matchShape and survive filterOnTopic. Neither declares a Signal, so an
empty composer still reports what the organization needs rather than proposing
paperwork.

Chip text is the coupling with nothing else holding it together: no page
context declares `capabilities`, so every server suggestion dispatches as its
own TEXT and is answered by whichever skill's trigger that text matches. A
renamed chip would open nothing, silently. Asserted on the frontend side.

The down migration drops `employee_role_status` and keeps `english_level`,
which is shared with job_postings.english_required and
job_applications.english_level. Rolled back and re-applied against the
database to prove it, not asserted.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PJvibeSc1JYXjatankqM1g
2026-09-02 15:29:25 +05:30

164 lines
8.8 KiB
Python
Executable File

#!/usr/bin/env python3
"""Emit go-api/internal/domain/resources_gen.go from the live PostgreSQL schema.
Column names, types, enum values and nullability are read out of
information_schema so they can never drift from the migrations. The per-resource
metadata below (path, default sort, default limit, operations, required fields)
comes from docs/api-contract.md and is the only hand-maintained part.
Usage: make gen-resources
"""
import collections
import json
import os
import subprocess
import sys
DB = os.environ.get("DATABASE_NAME", "Krow-force")
HOST = os.environ.get("DATABASE_HOST", "127.0.0.1")
PORT = os.environ.get("DATABASE_PORT", "5432")
USER = os.environ.get("DATABASE_USER", "postgres")
def q(sql):
out = subprocess.run(
["psql", "-h", HOST, "-p", PORT, "-U", USER, "-d", DB, "-tAF", "\x1f", "-c", sql],
capture_output=True, text=True)
if out.returncode:
sys.exit(out.stderr)
return [l.split("\x1f") for l in out.stdout.strip().split("\n") if l.strip()]
# ops='' means the resource has a table and is seeded, but serves no HTTP
# endpoint (see api-contract.md §2). It still needs a descriptor so the seeder
# can write it.
META = {
'job_postings': dict(name='JobPosting', path='job-postings', sort='-created_date', limit=100, ops='List|Get|Create|Update', req=['title']),
'job_applications':dict(name='JobApplication',path='job-applications',sort='-ai_score', limit=200, ops='List|Create|Update|Delete', req=['job_posting_id','applicant_name','email']),
'ai_interviews': dict(name='AIInterview', path='ai-interviews', sort='-created_date', limit=100, ops='List|Create', req=['application_id','job_posting_id']),
'staff': dict(name='Staff', path='staff', sort='-created_date', limit=100, ops='List|Create|Update', req=['name','email','hire_date']),
'worker_profiles': dict(name='WorkerProfile', path='worker-profiles', sort='-krow_score', limit=500, ops='List|Create|Update', req=['full_name','email']),
'courses': dict(name='Course', path='courses', sort='-created_date', limit=200, ops='List|Get|Create|Update', req=['title'], orgnull=True),
'learning_paths': dict(name='LearningPath', path='learning-paths', sort='-created_date', limit=100, ops='List', req=['name'], orgnull=True),
'role_categories': dict(name='RoleCategory', path='role-categories', sort='-created_date', limit=100, ops='List|Create', req=['name']),
'certifications': dict(name='Certification', path='certifications', sort='-created_date', limit=200, ops='List|Create|Delete', req=['name']),
'user_activity': dict(name='UserActivity', path='user-activity', sort='-created_date', limit=500, ops='List|Create', req=['event_type']),
'evidence': dict(name='Evidence', path='evidence', sort='-created_date', limit=200, ops='List|Create|Update', req=['type','worker_email']),
'assignments': dict(name='Assignment', path='assignments', sort='-created_date', limit=500, ops='List|Create', req=['job_posting_id','worker_email','starts_at']),
'shift_records': dict(name='ShiftRecord', path='shift-records', sort='-created_date', limit=500, ops='List', req=[]),
'employee_roles': dict(name='EmployeeRole', path='employee-roles', sort='-created_date', limit=200, ops='List|Get|Create|Update', req=['worker_email','role_category']),
'badges': dict(name='Badge', path='badges', sort='-created_date', limit=200, ops='', req=['name']),
}
ORDER = list(META)
# Columns the API never accepts from a request body, on every table that has
# them. The row's identity, its tenant and its timestamps are the server's.
READONLY = {'id', 'org_id', 'created_date', 'updated_date', 'legacy_id'}
# Columns that are server-owned on ONE table only.
#
# These name a *person*, and before Phase 3D a client could set them freely —
# which meant any authorization rule written on top of them could be defeated by
# the same request the rule was meant to constrain. A caller could reassign a
# worker profile to somebody else, or write an audit-log entry attributed to
# anyone in the organization.
#
# They are read-only here and filled in from the authenticated session instead;
# see the Derived rules in internal/domain/policy.go for which value each one
# receives and when.
SERVER_OWNED = {
'worker_profiles': {'user_id'},
'user_activity': {'user_id', 'user_email', 'user_name', 'account_type'},
'job_postings': {'created_by'},
'employee_roles': {'created_by'},
}
def kind(dt, udt, enums):
"""Classify a column.
USER-DEFINED covers both enums and extension types: citext reports as
USER-DEFINED too. Enum-ness is decided by whether pg_enum actually has
labels for the type, not by data_type alone — classifying citext as an
enum with no permitted values rejects every email the API is sent.
"""
if udt == 'uuid': return 'KindUUID', 'uuid'
if dt == 'ARRAY': return 'KindTextArray', 'text[]'
if dt == 'jsonb': return 'KindJSON', 'jsonb'
if dt in ('integer', 'smallint'): return 'KindInt', 'int'
if dt == 'bigint': return 'KindInt', 'bigint'
if dt == 'numeric': return 'KindFloat', 'numeric'
if dt == 'boolean': return 'KindBool', 'boolean'
if dt == 'timestamp with time zone': return 'KindTimestamp', 'timestamptz'
if dt == 'date': return 'KindDate', 'date'
if dt == 'USER-DEFINED' and enums.get(udt):
return 'KindEnum', udt
if udt == 'citext': return 'KindString', 'citext'
return 'KindString', 'text'
def main():
cols = q("""SELECT table_name, column_name, data_type, udt_name, is_nullable
FROM information_schema.columns
WHERE table_schema='public' AND table_name<>'schema_migrations'
ORDER BY table_name, ordinal_position""")
enums = collections.defaultdict(list)
for t, v in q("""SELECT t.typname, e.enumlabel FROM pg_type t
JOIN pg_enum e ON e.enumtypid=t.oid
JOIN pg_namespace n ON n.oid=t.typnamespace
WHERE n.nspname='public' ORDER BY t.typname, e.enumsortorder"""):
enums[t].append(v)
by_table = collections.defaultdict(list)
for t, c, dt, udt, nul in cols:
by_table[t].append((c, dt, udt, nul == 'YES'))
o = []
o.append('// Code generated by scripts/gen_resources.py. DO NOT EDIT BY HAND.')
o.append('// Regenerate with: make gen-resources')
o.append('//')
o.append('// Column names, types, enum values and nullability are read out of')
o.append('// information_schema so they cannot drift from the migrations. The')
o.append('// per-resource metadata (path, default sort, default limit, supported')
o.append('// operations, required fields) comes from docs/api-contract.md.')
o.append('')
o.append('package domain')
o.append('')
o.append('// AllResources is every resource the API serves.')
o.append('var AllResources = []*Resource{')
for tbl in ORDER:
m = META[tbl]
if m['ops']:
ops = ' | '.join('Op' + x for x in m['ops'].split('|'))
else:
# No operations means no routes are registered for this resource.
o.append(f'\t// {m["name"]} serves NO endpoint: useBadges has zero consumers and every')
o.append('\t// badge the UI renders comes from worker_profiles.earned_badges. The')
o.append('\t// descriptor exists so the seeder can write the table. api-contract.md §2.')
ops = '0'
o.append('\t{')
o.append(f'\t\tName: {json.dumps(m["name"])}, Path: {json.dumps(m["path"])}, Table: {json.dumps(tbl)},')
o.append(f'\t\tDefaultSort: {json.dumps(m["sort"])}, DefaultLimit: {m["limit"]},')
o.append(f'\t\tOps: {ops},')
if m.get('orgnull'):
o.append('\t\tOrgNullable: true,')
o.append('\t\tColumns: []Column{')
for c, dt, udt, nullable in by_table[tbl]:
k, cast = kind(dt, udt, enums)
p = [f'Name: {json.dumps(c)}', f'Kind: {k}', f'PGType: {json.dumps(cast)}']
if not nullable: p.append('NotNull: true')
if c in READONLY or c in SERVER_OWNED.get(tbl, ()):
p.append('ReadOnly: true')
if c in m['req']: p.append('Required: true')
if k == 'KindEnum':
p.append('Enum: []string{' + ', '.join(json.dumps(v) for v in enums[udt]) + '}')
o.append('\t\t\t{' + ', '.join(p) + '},')
o.append('\t\t},')
o.append('\t},')
o.append('}')
sys.stdout.write('\n'.join(o) + '\n')
if __name__ == '__main__':
main()