AI and Knowledge Graph Engineer
Location: Manchester, UKContract Position
About the Project
Join a specialist software development team delivering a new, business-critical technology platform for an established international organisation. This greenfield development project involves modern cloud architecture, data-intensive applications, and AI-enabled capabilities. You will work as part of a multidisciplinary team alongside experienced software, AI, infrastructure, and design professionals, taking the platform from development through to production. The project is delivered for a live business environment, requiring a strong focus on security, scalability, maintainability, data protection, and production readiness.
Role Purpose
As the AI and Knowledge Graph Engineer, you will build production Python services for information extraction, entity resolution, semantic search, and knowledge graphs. This role combines applied machine learning with semantic-data engineering. Applicants must have deployed RDF systems and production AI services, with tested code, reproducible evaluation, and practical fault diagnosis.
Main Responsibilities
- Knowledge Graph Architecture: Design RDF models that retain source evidence, temporal context, and provenance. Test graph-store behaviour against representative data and queries, and document modelling assumptions and trade-offs.
- Ontology, Taxonomy, and Standards: Work with RDF, SPARQL, SHACL, PROV-O, SKOS, JSON-LD, and bounded Web Ontology Language reasoning. Define ontology modules, namespaces, and stable identifiers. Manage named graphs, vocabulary mappings, versioning, and deprecation, retaining mapping provenance and tests for changes.
- Graph Projection and Validation: Build repeatable RDF ingestion and mapping processes, including incremental updates and full rebuilds. Validate changes before publication, optimise SPARQL queries, and maintain provenance, tenant separation, and access checks. Integrate graphs with relational systems and event-driven services.
- Document and Information Extraction: Develop Python and FastAPI services to extract structured information from documents, forms, messages, and transcripts. Use schema validation and preserve source locations, model versions, confidence, and extraction errors. Evaluate extraction against labelled examples and ensure unsupported output remains visible for validation.
- Entity Resolution: Develop record-linkage methods combining reliable identifiers with probabilistic comparison. Evaluate linkage quality, preserve source identifiers, and support reversible merge decisions. Test for missed matches and incorrect links, and assess their impact on downstream data quality.
- Search and Matching: Implement and evaluate lexical, vector, and hybrid retrieval, including taxonomy expansion and graph queries. Work with embedding selection, text segmentation, metadata filters, and reranking. Ensure scoring is reproducible and explanations are grounded in evidence.
- Model Evaluation: Create labelled datasets and repeatable tests for extraction, entity resolution, retrieval, ranking, and explanation quality. Build Python and SQL transformations for versioned datasets, prevent data leakage, and retain source and model versions. Monitor model quality after release and recommend retraining or rollback as needed.
- Responsible Artificial Intelligence: Apply data minimisation and access controls, test for prompt injection and unintended disclosure, and ensure sensitive attributes are handled appropriately. Implement controlled fallback, suspension, and rollback, and check model and dataset licences before adoption.
Essential Technical Experience
- At least six years of professional experience in artificial intelligence, machine learning, or semantic-data engineering.
- Demonstrated responsibility for production AI services and RDF knowledge graphs, including implementation, testing, and operational troubleshooting.
- Strong Python, FastAPI, and Pydantic experience, with skills in SQL, data transformation, and automated testing.
- Practical experience in information extraction, natural-language processing, embeddings, hybrid search, and reranking, with reproducible evaluation on labelled data.
- Production RDF and SPARQL experience, including ontology design, SHACL validation, provenance, taxonomy modelling, and query optimisation.
- Experience with RDF-capable graph stores such as GraphDB, Stardog, Apache Jena, or Amazon Neptune.
- Experience integrating model services with transactional applications, handling asynchronous data processing, and maintaining versioned datasets.
- Understanding of data privacy, model security, and prompt-injection risks.
- Ability to define service contracts, implement bounded execution, timeouts, retries, and resource limits for model workloads.
- Integration testing for malformed output, provider failure, and safe fallback.
Desirable Experience
- Experience with Microsoft Foundry, Azure Machine Learning, MLflow, Azure AI Search, or OpenSearch.
- Large-scale document processing, model gateways, learning-to-rank methods, and temporal or provenance-aware graphs.
- Familiarity with Parquet or Delta datasets, lakehouse processing, PyTorch, sentence-transformers, spaCy, or scikit-learn.
- Experience with practical fairness or subgroup evaluation.