# Training Data Related Articles

HTX News Center provides the latest articles and in-depth analysis on "Training Data", covering market trends, project updates, tech developments, and regulatory policies in the crypto industry.

You Use Claude and Codex Every Day, but Meta Has Restricted Internal Use

In May, Meta imposed internal restrictions on its engineers regarding the use of Claude Code and Codex, two widely used AI programming tools. Despite being a major client, Meta's guidelines, still in effect, prohibit these external models from being used for specific tasks to prevent potential "escalations with partners." The core concern is "distillation"—the risk that outputs from Claude or Codex could inadvertently contaminate the training data and evaluation processes for Meta's in-house AI coding assistant, MetaCode. If MetaCode is trained or evaluated using data generated by these external models, it risks learning their capabilities rather than developing its own, blurring the line of intellectual origin. The restrictions are precise: engineers cannot use the external models to generate test questions, debug source code, or suggest test cases. AI-generated content is also barred from environments accessible to MetaCode. However, AI can still assist with peripheral tasks like workflow setup and code organization, provided all outputs are manually reviewed. This caution reflects a broader industry dilemma. While distillation is a common technique, using a competitor's model output for training raises legal and ethical questions about the ownership of derived capabilities. Contractual terms from companies like OpenAI and Anthropic explicitly forbid using their outputs to build competing products, putting enforcement power in the hands of rivals. The move is also financially motivated, as Meta seeks to reduce its hefty internal AI spending, estimated in the billions this year. Meta's policy illustrates the delicate balance companies must strike: leveraging powerful external AI tools while safeguarding the integrity and independence of their own AI development. As AI systems increasingly help build other AIs, distinguishing the origin of capabilities becomes a fundamental challenge for the entire industry.

marsbit06/30 13:13

You Use Claude and Codex Every Day, but Meta Has Restricted Internal Use

marsbit06/30 13:13

Why Sam Altman's 'Water and Electricity Theory' Sparks Copyright Controversy

OpenAI CEO Sam Altman's recent statement that "intelligence will become a utility like electricity or water" has sparked significant controversy, primarily around copyright issues and the nature of AI development. While positioning AI as a utility serves as a compelling narrative for infrastructure investors, critics argue the analogy is flawed in three key areas. First, there's a fundamental "property gap." Traditional utilities like water and power create new, physical infrastructure from scratch. In contrast, major AI models are trained by reorganizing vast amounts of existing human-created content—books, articles, code, etc.—often scraped from the web without explicit permission or compensation to creators. This "free acquisition, paid resale" model is seen by many as ethically problematic. Second, there's a "pricing gap." True public utilities are typically regulated to ensure universal service with non-discriminatory, cost-plus pricing. AI's token-based pricing, however, involves significant price discrimination (e.g., output tokens costing much more than input tokens) and is designed for revenue maximization, not equitable access. Third, a "governance gap" exists. Utilities operate under public oversight, while AI pricing and development are currently controlled by a few private companies. Furthermore, the industry's own shift toward buying licensed training data (e.g., deals with Reddit or news publishers) undermines its previous legal reliance on "fair use" for freely scraped data. In conclusion, while AI is indeed becoming a foundational technology, calling it a public utility remains contentious. The title requires not just scale and a pay-per-use model, but also credible solutions for data provenance, equitable pricing, and public governance.

marsbit05/27 10:03

Why Sam Altman's 'Water and Electricity Theory' Sparks Copyright Controversy

marsbit05/27 10:03

活动图片