POLICYIQ: AN AI-POWERED POLICY CONFLICT VERIFICATION SYSTEM | IJCSE Volume 10 – Issue 5 | IJCSE-V10I5P24

IJCSE International Journal of Computer Science Engineering Logo

International Journal of Computer Science Engineering Techniques

ISSN: 2455-135X
Volume 10, Issue 5  |  Published:
Author

Abstract

Organisations, governments and regulated industries operate under large, continuously revised bodies of policy. As these documents evolve across versions, editions and jurisdictions, overlapping, contradictory or silently modified clauses accumulate and create compliance, governance and legal risk that manual review cannot scale to detect. To resolve these systemic bottlenecks, this paper presents PolicyIQ, an enterprise-grade artificial intelligence platform for automated policy conflict verification and version-aware change intelligence. PolicyIQ ingests heterogeneous policy documents through a robust extraction pipeline with optical character recognition (OCR) fallback, decomposes them into atomic clauses using a hybrid rule-based and semantic segmentation engine, and encodes each clause into a 384-dimensional dense vector using a Sentence-BERT transformer (MiniLM-L6). A cross-version cosine-similarity matrix drives a multi-stage classifier that labels every clause transition as duplicate, modified, new, removed or direct-conflict, augmented by deontic, numeric and negation-scope heuristics for typed contradiction detection. A distribution-aware adaptive thresholding scheme replaces brittle fixed cut-offs by deriving decision boundaries from the per-corpus similarity distribution. Built on a decoupled four-tier microservices architecture utilising a FastAPI security gateway, asynchronous Celery workers, PostgreSQL relational storage and a Qdrant/FAISS vector store, PolicyIQ delivers explainable dual-audience summaries while maintaining data privacy through JWT authentication and role-based access control. On a controlled synthetic policy corpus the platform attains target performance of 91.6% precision for adaptive conflict classification and 0.944 segmentation F1, with sub-second clause-level inference, demonstrating a reproducible, provenance-aware pathway from raw policy text to actionable conflict intelligence.

Keywords

PolicyIQ, Artificial Intelligence, Policy Conflict Verification, Semantic Textual Similarity, Sentence-BERT, Clause Segmentation, Adaptive Thresholding, Contradiction Detection, Vector Search, Microservices Architecture, Legal NLP, Version Diffing, Explainable Summarisation.

Conclusion

This paper presented PolicyIQ, a version-aware, multi-stage artificial intelligence platform for automated policy conflict verification and change intelligence. By unifying robust text extraction, hybrid clause segmentation, dense Sentence-BERT embeddings, cross-version typed conflict classification, distribution-aware adaptive thresholding and dual-audience summarisation within a hardened four-tier microservices architecture, PolicyIQ demonstrates a reproducible pathway from raw policy text to explainable, actionable conflict intelligence. The design directly targets the gaps left by lexical diffing and version-blind similarity tools, replacing opaque scalar scores with typed, provenance-preserving verdicts secured behind an authenticated, role-based API. On a controlled synthetic benchmark the platform meets its design targets for segmentation and adaptive conflict classification while sustaining sub-second clause-level inference, establishing a solid foundation for rigorous, human-annotated evaluation.

References

[1] T. Mikolov, K. Chen, G. Corrado, and J. Dean, β€œEfficient estimation of word representations in vector space,” in Proc. ICLR Workshop, 2013. [2] A. Vaswani et al., β€œAttention is all you need,” in Proc. NeurIPS, Long Beach, CA, USA, 2017, pp. 5998–6008. [3] J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, β€œBERT: Pre-training of deep bidirectional transformers for language understanding,” in Proc. NAACL-HLT, 2019, pp. 4171–4186. [4] N. Reimers and I. Gurevych, β€œSentence-BERT: Sentence embeddings using Siamese BERT-networks,” in Proc. EMNLP-IJCNLP, 2019, pp. 3982–3992. [5] S. R. Bowman, G. Angeli, C. Potts, and C. D. Manning, β€œA large annotated corpus for learning natural language inference,” in Proc. EMNLP, 2015, pp. 632–642. [6] Y. Koreeda and C. D. Manning, β€œContractNLI: A dataset for document-level natural language inference for contracts,” in Findings of EMNLP, 2021, pp. 1907–1919. [7] LegalWiz, β€œA multi-agent generation framework for contradiction detection in legal documents,” arXiv preprint arXiv:2510.03418, 2025. [8] β€œNLP-based regulatory compliance: Using GPT-4.0 to decode regulatory documents,” arXiv preprint arXiv:2412.20602, 2024. [9] M. Lippi et al., β€œDeep learning for conflicting statements detection in text,” PeerJ Preprints, vol. 6, e26589v1, 2018. [10] S. Gupta et al., β€œIdentifying contradictions in the legal proceedings using natural language models,” SN Computer Science, vol. 3, no. 3, 2022. [11] A. Mourad and H. Jebbaoui, β€œConflict detection and resolution of XACML policies,” in Proc. IEEE Symp. Computers and Communications, 2014. [12] M. A. Rahman et al., β€œClustering-based approach for anomaly detection in XACML policies,” in Proc. SECRYPT, 2017, pp. 548–553. [13] H. Shah and P. Kumar, β€œAutomated interpretation of regulations using NLP for compliance-centric analysis of legal texts,” ResearchGate preprint, 2024. [14] R. Mohan et al., β€œDetecting compliance of privacy policies with data protection laws,” arXiv preprint arXiv:2102.12362, 2021. [15] D. Silva et al., β€œAnalysing similarities between legal court documents using NLP approaches based on transformers,” PLoS ONE / PMC, 2024. [16] I. Chalkidis et al., β€œLEGAL-BERT: The muppets straight out of law school,” in Findings of EMNLP, 2020, pp. 2898–2904. [17] J. Johnson, M. Douze, and H. JΓ©gou, β€œBillion-scale similarity search with GPUs,” IEEE Trans. Big Data, vol. 7, no. 3, pp. 535–547, 2021. [18] Yu. A. Malkov and D. A. Yashunin, β€œEfficient and robust approximate nearest neighbor search using hierarchical navigable small world graphs,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 42, no. 4, pp. 824–836, 2020. [19] R. Mihalcea and P. Tarau, β€œTextRank: Bringing order into texts,” in Proc. EMNLP, 2004, pp. 404–411. [20] R. S. Sandhu, E. J. Coyne, H. L. Feinstein, and C. E. Youman, β€œRole-based access control models,” IEEE Computer, vol. 29, no. 2, pp. 38–47, 1996. [21] C. D. Manning, P. Raghavan, and H. SchΓΌtze, Introduction to Information Retrieval. Cambridge, U.K.: Cambridge Univ. Press, 2008. [22] R. Smith, β€œAn overview of the Tesseract OCR engine,” in Proc. ICDAR, 2007, pp. 629–633. [23] Y. Liu et al., β€œRoBERTa: A robustly optimized BERT pretraining approach,” arXiv preprint arXiv:1907.11692, 2019. [24] W. Wang et al., β€œMiniLM: Deep self-attention distillation for task-agnostic compression of pre-trained transformers,” in Proc. NeurIPS, 2020. [25] S. RamΓ­rez, β€œFastAPI: Modern, fast web framework for building APIs with Python,” software documentation, 2023. [26] Qdrant Team, β€œQdrant: Vector similarity search engine,” software documentation, 2024.
Β© 2025 International Journal of Computer Science Engineering Techniques (IJCSE).

Related Post

Submit Your Paper