Dev.to · 6 min read

Catching Cross-Language Copy-Paste Debt with Static Analysis and AI Pair-Programmers

Catching Cross-Language Copy-Paste Debt with Static Analysis and AI Pair-Programmers

Originally published on tamiz.pro. Introduction Most linters check within a single language or project boundary. They miss copy-paste debt that spans repositories, languages, or even teams. This is the story of how I combined static analysis with AI pair-programmers to catch real copy-paste debt across 220 languages — and why traditional tooling fell short. The Problem: Linters Are Local Linters excel at syntax and local style enforcement. But copy-paste debt often crosses project boundaries: Duplicated utility functions in JavaScript and TypeScript Cloned database access patterns in Go and Python Repetitive configuration blocks in YAML, JSON, and TOML Copied authentication logic between services written in different languages Traditional static analysis tools like ESLint, Pyflakes, or golangci-lint operate within a single language ecosystem. Even cross-cutting tools like SonarQube struggle when duplication spans truly disparate codebases. The Approach: Multi-Language Hashing Step 1: Normalize Source Code Before comparing code, we needed to normalize it: import re def normalize_code(code: str) -> str: # Remove comments code = re.sub(r'#.*', '', code) code = re.sub(r'//.*', '', code) code = re.sub(r'/\*.*?\*/', '', code, flags=re.DOTALL) # Normalize whitespace code = re.sub(r'\s+', ' ', code) # Normalize strings code = re.sub(r'"[^"]*"', '""', code) code = re.sub(r"'[^']*'", "''", code) return code.strip() This normalization removes superficial differences while preserving structural similarity. Step 2: Generate Structural Hashes We used a combination of techniques: Token-based hashing: Convert normalized code to tokens and hash sequences AST-based comparison: For languages with available parsers, compare abstract syntax trees N-gram analysis: Break code into overlapping n-grams (n=5 worked well) from collections import defaultdict class MultiLanguageDupDetector: def __init__(self): self.hashes = defaultdict(list) self.min_hash_length = 50 # tokens def add_file(self, file_path: str, language: str, code: str): normalized = normalize_code(code) tokens = self.tokenize(normalized, language) if len(tokens) < self.min_hash_length: return # Generate rolling hashes for i in range(len(tokens) - self.min_hash_length + 1): chunk = tokens[i:i + self.min_hash_length] hash_key = hash(tuple(chunk)) self.hashes[hash_key].append({ 'file': file_path, 'language': language, 'start_line': self.get_line_number(code, i), 'content': ' '.join(chunk) }) def find_duplicates(self, min_matches=3): duplicates = [] for hash_key, locations in self.hashes.items(): if len(locations) >= min_matches: duplicates.append(locations) return duplicates Integrating AI Pair-Programmers Why AI? Static analysis catches exact and near-exact copies. But human developers don't copy-paste exact code — they tweak variable names, adjust formatting, and modify logic slightly. This is where AI pair-programmers shine. The Workflow Static analysis pre-filter: Identify candidate duplicates using hashing AI semantic comparison: Use AI to assess whether candidates are truly duplicated logic Confidence scoring: Rank findings by likelihood of actual duplication import openai class AISimilarityChecker: def __init__(self, api_key: str): self.client = openai.OpenAI(api_key=api_key) def check_similarity(self, code_snippet_1: str, code_snippet_2: str) -> dict: prompt = f""" Compare these two code snippets. Are they duplicates that should be refactored? Rate on a scale of 1-10 how similar they are in logic and structure. Snippet 1: {code_snippet_1} Snippet 2: {code_snippet_2} Respond with JSON: {{"score": , "reasoning": ""}} """ response = self.client.chat.completions.create( model="gpt-4-turbo-preview", messages=[{"role": "user", "content": prompt}], response_format={"type": "json_object"} ) return json.loads(response.choices[0].message.content) Scaling Across 220 Languages Language-Specific Parsers Not all languages have mature parsing libraries. For less common languages, we used: Tree-sitter: Supports 50+ languages with incremental parsing ANTLR grammars: For languages without tree-sitter support Fallback tokenization: Regex-based tokenizers for unsupported languages Performance Considerations Processing thousands of files across 220 languages required optimization: Parallel processing: Used multiprocessing pools Incremental updates: Only reprocess changed files Bloom filters: Quickly eliminate non-matching candidates from concurrent.futures import ProcessPoolExecutor import multiprocessing def process_repository_batch(repos_batch): detector = MultiLanguageDupDetector() for repo in repos_batch: # Process each file in repository pass return detector.hashes def scan_all_repositories(repositories, num_workers=None): if num_workers is None: num_workers = multiprocessing.cpu_count() - 2 batch_size = max(1, len(repositories) // num_workers) batches = [ repositories[i:i + batch_size] for i in range(0, len(repositories), batch_size) ] with ProcessPoolExecutor(max_workers=num_workers) as executor: results = list(executor.map(process_repository_batch, batches)) # Merge results merged_hashes = defaultdict(list) for result in results: for hash_key, locations in result.items(): merged_hashes[hash_key].extend(locations) return merged_hashes Real Findings After scanning 500+ repositories, we found several categories of copy-paste debt: Category 1: Cross-Language Utility Duplication // JavaScript utility function formatDate(date) { const year = date.getFullYear(); const month = String(date.getMonth() + 1).padStart(2, '0'); const day = String(date.getDate()).padStart(2, '0'); return `${year}-${month}-${day}`; } // TypeScript clone export function formatDate(inputDate: Date): string { const year = inputDate.getFullYear(); const month = String(inputDate.getMonth() + 1).padStart(2, '0'); const day = String(inputDate.getDate()).padStart(2, '0'); return `${year}-${month}-${day}`; } Category 2: Configuration Template Duplication Nearly identical Docker Compose files across different environments: # development.yaml version: '3.8' services: app: build: . ports: - "3000:3000" environment: - NODE_ENV=development # staging.yaml version: '3.8' services: app: build: . ports: - "3000:3000" environment: - NODE_ENV=staging Why Linters Missed These Several reasons explain why traditional linters failed: Language isolation: Linters don't cross language boundaries Project scope: Linters operate on individual projects, not organization-wide Semantic understanding: Most linters lack deep semantic analysis Configuration overhead: Setting up cross-project analysis is complex Best Practices When to Use This Approach Large organizations with polyglot codebases Teams maintaining similar projects across languages Legacy systems with accumulated technical debt Implementation Tips Start small: Begin with high-value repositories Tune thresholds: Adjust similarity scores based on false positive tolerance Automate reviews: Integrate findings into CI/CD pipelines Track improvements: Monitor debt reduction over time Frequently Asked Questions Q: Do I need AI for copy-paste detection? A: Not always. For exact duplicates, static analysis suffices. AI helps with semantic similarities where code logic is reused but implementations differ slightly. Q: How do I handle false positives? A: Implement confidence scoring and manual review workflows. Start with high-confidence matches and gradually lower thresholds as you fine-tune the system. Q: What's the performance impact? A: With proper parallelization and incremental processing, scanning large codebases should complete within minutes. Use bloom filters and other probabilistic data structures to reduce computation. Conclusion Combining static analysis with AI creates a powerful system for detecting cross-language copy-paste debt that traditional linters miss. The key is normalization, multi-language support, and semantic understanding through AI assistance. While this approach requires more setup than standard linters, the technical debt reduction often justifies the investment. For teams managing complex, polyglot environments, this hybrid approach provides visibility into duplication patterns that would otherwise remain hidden.

This is a summary aggregated from Dev.to. Read the complete article on the original site:

Read full article at Dev.to

More AI & Machine Learning News